| soundgen | R Documentation |
Generates a sequence ("bout") of one or more vocalizations ("syllables") with pauses between them. Two basic components are synthesized: a periodic component (the sum of sine waves with frequencies that are multiples of the fundamental frequency) and an aperiodic noise component. Both components can be filtered with independently specified vocal tract transfer functions ("formants"). Intonation and amplitude contours can be applied both within each syllable and across multiple syllables. Suggested application: synthesis of animal calls and human nonverbal vocalizations (not speech). For more information, see https://cogsci.se/soundgen/sound_generation.html.
soundgen(
repeatBout = 1,
nSyl = 1,
sylLen = 500,
pauseLen = 200,
addSilence = 100,
ampl = NA,
amplGlobal = NA,
attackLen = 50,
pitch = c(100, 150, 100),
pitchGlobal = NA,
rolloff = -12,
rolloffOct = 0,
rolloffKHz = 0,
rolloffExact = NULL,
glottis = 0,
pitchFloor = 1,
pitchCeiling = 3500,
pitchSamplingRate = 16000,
noise = NULL,
rolloffNoise = -4,
noiseFlatSpec = 1200,
rolloffNoiseExp = 0,
formants = c(860, 1430, 2900, 4100),
formantsNoise = NA,
formantDep = 1,
formantDepStoch = 1,
formantWidth = 1,
formantCeiling = NULL,
formantLocking = 0,
vocalTract = NA,
mouth = NULL,
lipRad = 6,
noseRad = 4,
mouthOpenThres = 0,
amDep = 0,
amFreq = 30,
amType = c("logistic", "sine"),
amShape = 0,
vibratoFreq = 5,
vibratoDep = 0,
jitterDep = 0,
jitterLen = 1,
shimmerDep = 0,
shimmerLen = 1,
subRatio = 2,
subDep = 0,
nonlinBalance = 100,
nonlinRandomWalk = NULL,
shortestEpoch = 300,
temperature = 0.025,
tempEffects = list(),
maleFemale = 0,
creakyBreathy = 0,
plot = FALSE,
play = FALSE,
saveAudio = FALSE,
invalidArgAction = c("adjust", "abort", "ignore"),
smoothing = list(interpol = "splineFC", discontThres = 0.05, jumpThres = 0.01),
samplingRate = 16000,
windowLength = 50,
step = NULL,
overlap = 75,
wn = "gaussian",
dynamicRange = 80,
...
)
repeatBout |
number of times the whole bout should be repeated |
nSyl |
number of syllables in the bout; pitchGlobal, amplGlobal, and formants span multiple syllables, but not multiple bouts |
sylLen |
average duration of each syllable, ms (vectorized) |
pauseLen |
average duration of pauses between syllables, ms (can be
negative between bouts to overlap them: force with invalidArgAction =
'ignore') (vectorized). If there are multiple bouts, the first value of
|
addSilence |
silence before and after the bout, ms: a vector of length 1 (symmetric) or 2 (different duration of silence before/after the sound) |
ampl |
amplitude envelope (dB, 0 = max amplitude): a numeric vector or anchor format |
amplGlobal |
global amplitude envelope spanning multiple syllables (dB, 0 = no change) (anchor format) |
attackLen |
duration of fade-in / fade-out at each end of syllables and noise (ms): a vector of length 1 (symmetric) or 2 (separately for fade-in and fade-out) |
pitch |
fundamental frequency within one syllable (anchor format). NAs
in pitch vectors are accepted and are converted into voiceless fragments,
but then set |
pitchGlobal |
unlike |
rolloff |
the rate at which f0 harmonics in the spectrum become weaker,
dB/oct (anchor format for all rolloff-related parameters). More negative
rolloff = weaker upper harmonics; see |
rolloffOct, rolloffKHz |
rolloff may be constant throughout the spectrum, or it may vary with each octave above f0 (rolloffOct) or per kHz increase in f0 above the baseline of 200 Hz (rolloffKHz); for both parameters, positive values mean the rate of rolloff increases toward upper frequencies) |
rolloffExact |
user-specified relative amplitude of harmonics: a vector or matrix with one row per harmonic, scale 0 to 1 (overrides all other rolloff parameters) |
glottis |
duration of the closed phase of a glottal cycle (silent) in relation to the open phase, % (0 = no closed phase, 100 = closed phase as long as open phase, 200 = twice as long as the open phase, etc.); numeric vector or dataframe specifying time and value (anchor format). Use this effect sparingly: it is slow, and high values may affect harmonic composition and require a high sampling rate, especially in combination with high pitch |
pitchFloor, pitchCeiling |
lower & upper bounds of f0 |
pitchSamplingRate |
sampling frequency of the pitch contour only, Hz.
Low values reduce processing time. Set to |
noise |
intensity of turbulent noise (0 dB = same RMS as that of the periodic (voiced) component, negative values = less intense; anchor format). In soundgen 3.0, the noise component is always calibrated relative to the filtered harmonic component. When noise is present, the harmonic and noise components are filtered separately, their RMS amplitudes are normalized after filtering, and they are then mixed. Because noise can begin before the voiced part and continue after it, the time of noise anchors MUST be in ms, not [0, 1]; this is different from all other soundgen arguments that accept the anchor format with time either in ms or [0, 1] |
rolloffNoise, rolloffNoiseExp, noiseFlatSpec |
linear (rolloffNoise,
dB/kHz, anchor format) or exponential (rolloffNoiseExp, dB/oct, anchor
format) rolloff of the excitation source for the noise component (anchor
format) applied above |
formants |
a vector of formant frequencies (assuming formants are static
throughout the sound); a list of formant times, frequencies, amplitudes,
and bandwidths; or a character string referring to default presets for
speaker "M1" (implemented: "aoieu0"). NA or NULL means no formants, only
lip radiation (but a schwa is generated if |
formantsNoise |
the same as |
formantDep |
scale factor of formant amplitude (1 = no change relative
to amplitudes in |
formantDepStoch |
the amplitude of additional stochastic formants added above the highest specified formant, dB (only if temperature > 0) |
formantWidth |
scale factor of formant bandwidth (1 = no change) |
formantCeiling |
frequency to which stochastic formants are calculated to avoid losing energy in the upper part of the spectrum due to unmodeled resonances above the Nyquist frequencies, specified in multiples of Nyquist. If NULL (default), a faster theoretical correction is used (may fail for unusual sounds) |
formantLocking |
the approximate proportion of sound in which one of the harmonics is locked to the nearest formant, 0 = none, 1 = the entire sound (anchor format). In multi-syllable sounds, formant locking is applied separately to each syllable using the corresponding portion of the bout-level formant filter |
vocalTract |
the length of vocal tract, cm. Used for calculating formant
dispersion (for adding extra formants) and formant transitions as the mouth
opens and closes. If |
mouth |
mouth opening (0 to 1, 0.5 = neutral, i.e. no modification) (anchor format) |
lipRad |
the effect of lip radiation on source spectrum, dB/oct (the default of +6 dB/oct produces a high-frequency boost when the mouth is open) |
noseRad |
the effect of radiation through the nose on source spectrum,
dB/oct (the alternative to |
mouthOpenThres |
open the lips (switch from nose radiation to lip
radiation) when the mouth is open |
amDep |
amplitude modulation (AM) depth, %. 0: no change; 100: AM with amplitude range equal to the dynamic range of the sound (anchor format) |
amFreq |
AM frequency, Hz (anchor format) |
amType |
"logistic" = logistic (default), "sine" = sinusoidal |
amShape |
ignored if amType = "sine", otherwise determines the shape of non-sinusoidal AM: 0 = ~sine, -1 = notches, +1 = clicks (anchor format) |
vibratoFreq |
the rate of regular pitch modulation, or vibrato, Hz (anchor format) |
vibratoDep |
the depth of vibrato, semitones (anchor format) |
jitterDep |
cycle-to-cycle random pitch variation, semitones (anchor format) |
jitterLen |
duration of stable periods between pitch jumps, ms. Use a low value for harsh noise, a high value for irregular vibrato or shaky voice (anchor format) |
shimmerDep |
random variation in amplitude between individual glottal cycles (0 to 100% of original amplitude of each cycle) (anchor format) |
shimmerLen |
duration of stable periods between amplitude jumps, ms. Use a low value for harsh noise, a high value for shaky voice (anchor format) |
subRatio |
a positive integer giving the ratio of f0 (the main fundamental) to g0 (a lower frequency): 1 = no subharmonics, 2 = period doubling regardless of pitch changes, 3 = period tripling, etc. |
subDep |
the depth of subharmonics relative to the main frequency component (f0), %. 0: no subharmonics; 100: g0 harmonics are as strong as the nearest f0 harmonic (anchor format) |
nonlinBalance |
hyperparameter for regulating the (approximate) proportion of sound with different regimes of pitch effects (none / subharmonics only / subharmonics and jitter). 0% = no nonlinear phenomena; 100% = the entire sound has jitter + subharmonics. Ignored if temperature = 0 |
nonlinRandomWalk |
a numeric vector specifying the timing of nonlinear regimes: 0 = none, 1 = subharmonics, 2 = subharmonics + jitter + shimmer |
shortestEpoch |
minimum duration of each epoch with unchanging subharmonics regime or formant locking, in ms |
temperature |
hyperparameter for regulating the amount of stochasticity in sound generation |
tempEffects |
a list of scaling coefficients regulating the effect of temperature on particular parameters. To change, specify just those pars that you want to modify (1 = default, 0 = no stochastic behavior).
|
maleFemale |
hyperparameter for shifting f0 contour, formants, and vocalTract to make the speaker appear more male (-1...0) or more female (0...+1); 0 = no change |
creakyBreathy |
hyperparameter for an adjustment of voice quality from creaky (-1) to breathy (+1); 0 = no change |
plot |
if TRUE, plots a spectrogram |
play |
if TRUE, plays the synthesized sound using the default player on
your system. If character, passed to |
saveAudio |
if TRUE, saves the result as "soundgen.wav" in the working directory |
invalidArgAction |
what to do if an argument is invalid or outside the permitted range: 'adjust' = reset to default value, 'abort' = stop execution, 'ignore' = throw a warning and continue (may crash) |
smoothing |
a list of parameters passed to |
samplingRate |
sampling rate of the output (Hz) |
windowLength |
length of the FFT window (ms) |
step |
step between successive windows (ms); if provided, overrides
|
overlap |
overlap between successive windows (0–100%) |
wn |
wn window type accepted by |
dynamicRange |
dynamic range (dB). Harmonics and noise more than dynamicRange under maximum amplitude are discarded to save computational resources |
... |
other plotting parameters passed to |
The synthesized waveform as a numeric vector normalized to [-1, 1]
and sampled at the specified samplingRate. Note that the sampling
rate may be increased internally in case the pitch is high, with a warning.
repeatBout, nSyl, sylLen, pauseLen, addSilence
ampl, amplGlobal, attackLen
pitch, pitchGlobal, rolloff, rolloffOct, rolloffKHz, rolloffExact, glottis, pitchFloor, pitchCeiling, pitchSamplingRate
noise, rolloffNoise, rolloffNoiseExp, noiseFlatSpec
formants, formantsNoise, formantDep, formantDepStoch, formantWidth, formantCeiling, formantLocking, vocalTract, lipRad, noseRad, mouth, mouthOpenThres
amDep, amFreq, amType, amShape, vibratoFreq, vibratoDep
jitterDep, jitterLen, shimmerDep, shimmerLen, subRatio, subDep, nonlinBalance, nonlinRandomWalk, shortestEpoch
temperature, tempEffects, maleFemale, creakyBreathy
plot, play, saveAudio, ...
smoothing, invalidArgAction, samplingRate, windowLength, step, overlap, wn, dynamicRange
soundgen() and some other functions in this package accept arguments that
define a time series (e.g., pitch or amplitude contour). These arguments can
be provided as either numeric vectors or a list of anchors - points through
which the contour is interpolated. This "anchor format" can be a dataframe or
list with two elements: $time (ms or 0 to 1) and value of each
anchor. Ex.: soundgen(pitch = list(time = c(0, 0.2, 1),
value = c(310, 340, 280))).
getFormantFilter beat
# Detailed documentation: https://cogsci.se/soundgen/sound_generation.html
# A gallery of examples with code: https://cogsci.se/soundgen/demos.html
# A GUI for soundgen is available as a Shiny app.
# Type "soundgen_app()" to open it in your default browser
# Set "playback" to TRUE for default system player or the name of preferred
# player (eg "aplay") to play back the audio from examples
playback = FALSE # or TRUE, 'aplay', 'vlc', etc. (see ?playme)
sound = soundgen(play = playback)
# spectrogram(sound, 16000)
# playme(sound)
# Control of intonation, amplitude envelope, formants
s0 = soundgen(
pitch = c(300, 390, 250),
ampl = data.frame(time = c(0, 50, 300), value = c(-5, -10, 0)),
attackLen = c(10, 50),
formants = c(600, 900, 2200),
play = playback
)
# Use the in-built collection of presets:
# names(presets) # speakers
# names(presets$Chimpanzee) # calls per speaker
s1 = eval(parse(text = presets$Chimpanzee$Scream_conflict)) # screaming chimp
# playme(s1)
s2 = eval(parse(text = presets$F1$Scream)) # screaming woman
# playme(s2, 18320)
# presets of some vowels and consonants
names(presets$M1$Formants$vowels)
soundgen(sylLen = 500, formants = 'aoieu0', play = playback)
## Not run:
# Ultrasound - need to adjust some defaults:
soundgen(
sylLen = 10, # just 10 ms
attackLen = 1, # should be very short for short vocalizations
addSilence = 2,
pitch = c(45000, 35000, 65000, 60000), # 35-60 kHz
rolloff = -12,
rolloffKHz = 0,
formants = NA, # no formants (or set vocal tract length)
samplingRate = 350000, # at least ~10 times the max f0
pitchSamplingRate = 350000, # the same as samplingRate
windowLength = .25, # need very short window lengths for USV
pitchCeiling = 90000, # max allowed pitch
invalidArgAction = 'ignore', # override the ranges allowed by default
temperature = 1e-4,
plot = TRUE
)
# soundgen plays Bach
dt = otherToHz(
c('E4', 'D4', 'E4', 'C4', 'E4', 'B3', 'E4', 'A3', 'E4', 'G#3', 'E4',
'A3', 'E4', 'B3', 'E4', 'C4', 'E4', 'E3', 'E4', 'F#3', 'E4', 'G#3',
'E4', 'A3', 'E4', 'G#3', 'E4', 'A3', 'E4', 'B3', 'E4', 'C4'), 'notes')
out = numeric(0)
for (s in 1:length(dt)) {
syl_s = soundgen(sylLen = 100, pitch = dt[s], rolloff = -15,
formants = c(750, 1400, 2900, 3800), noise = -45,
attackLen = 50, addSilence = 0, temperature = .01)
syl_s = fade(soundgen:::matchLengths(syl_s, 0.1 * 16000), samplingRate = 16000)
out = c(out, syl_s[1:(0.1 * 16000)])
}
spectrogram(out, 16000, yScale = 'ERB')
playme(out, 16000)
## End(Not run)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.