soundgen: Generate a sound

View source: R/soundgen.R

soundgenR Documentation

Generate a sound

Description

Generates a sequence ("bout") of one or more vocalizations ("syllables") with pauses between them. Two basic components are synthesized: a periodic component (the sum of sine waves with frequencies that are multiples of the fundamental frequency) and an aperiodic noise component. Both components can be filtered with independently specified vocal tract transfer functions ("formants"). Intonation and amplitude contours can be applied both within each syllable and across multiple syllables. Suggested application: synthesis of animal calls and human nonverbal vocalizations (not speech). For more information, see https://cogsci.se/soundgen/sound_generation.html.

Usage

soundgen(
  repeatBout = 1,
  nSyl = 1,
  sylLen = 500,
  pauseLen = 200,
  addSilence = 100,
  ampl = NA,
  amplGlobal = NA,
  attackLen = 50,
  pitch = c(100, 150, 100),
  pitchGlobal = NA,
  rolloff = -12,
  rolloffOct = 0,
  rolloffKHz = 0,
  rolloffExact = NULL,
  glottis = 0,
  pitchFloor = 1,
  pitchCeiling = 3500,
  pitchSamplingRate = 16000,
  noise = NULL,
  rolloffNoise = -4,
  noiseFlatSpec = 1200,
  rolloffNoiseExp = 0,
  formants = c(860, 1430, 2900, 4100),
  formantsNoise = NA,
  formantDep = 1,
  formantDepStoch = 1,
  formantWidth = 1,
  formantCeiling = NULL,
  formantLocking = 0,
  vocalTract = NA,
  mouth = NULL,
  lipRad = 6,
  noseRad = 4,
  mouthOpenThres = 0,
  amDep = 0,
  amFreq = 30,
  amType = c("logistic", "sine"),
  amShape = 0,
  vibratoFreq = 5,
  vibratoDep = 0,
  jitterDep = 0,
  jitterLen = 1,
  shimmerDep = 0,
  shimmerLen = 1,
  subRatio = 2,
  subDep = 0,
  nonlinBalance = 100,
  nonlinRandomWalk = NULL,
  shortestEpoch = 300,
  temperature = 0.025,
  tempEffects = list(),
  maleFemale = 0,
  creakyBreathy = 0,
  plot = FALSE,
  play = FALSE,
  saveAudio = FALSE,
  invalidArgAction = c("adjust", "abort", "ignore"),
  smoothing = list(interpol = "splineFC", discontThres = 0.05, jumpThres = 0.01),
  samplingRate = 16000,
  windowLength = 50,
  step = NULL,
  overlap = 75,
  wn = "gaussian",
  dynamicRange = 80,
  ...
)

Arguments

repeatBout

number of times the whole bout should be repeated

nSyl

number of syllables in the bout; pitchGlobal, amplGlobal, and formants span multiple syllables, but not multiple bouts

sylLen

average duration of each syllable, ms (vectorized)

pauseLen

average duration of pauses between syllables, ms (can be negative between bouts to overlap them: force with invalidArgAction = 'ignore') (vectorized). If there are multiple bouts, the first value of pauseLen is used at the pause between bouts

addSilence

silence before and after the bout, ms: a vector of length 1 (symmetric) or 2 (different duration of silence before/after the sound)

ampl

amplitude envelope (dB, 0 = max amplitude): a numeric vector or anchor format

amplGlobal

global amplitude envelope spanning multiple syllables (dB, 0 = no change) (anchor format)

attackLen

duration of fade-in / fade-out at each end of syllables and noise (ms): a vector of length 1 (symmetric) or 2 (separately for fade-in and fade-out)

pitch

fundamental frequency within one syllable (anchor format). NAs in pitch vectors are accepted and are converted into voiceless fragments, but then set nSyl = 1

pitchGlobal

unlike pitch, these anchors are used to create a smooth contour of average f0 across multiple syllables. The values are in semitones relative to the existing pitch, i.e. 0 = no change (anchor format)

rolloff

the rate at which f0 harmonics in the spectrum become weaker, dB/oct (anchor format for all rolloff-related parameters). More negative rolloff = weaker upper harmonics; see getRolloff for more details

rolloffOct, rolloffKHz

rolloff may be constant throughout the spectrum, or it may vary with each octave above f0 (rolloffOct) or per kHz increase in f0 above the baseline of 200 Hz (rolloffKHz); for both parameters, positive values mean the rate of rolloff increases toward upper frequencies)

rolloffExact

user-specified relative amplitude of harmonics: a vector or matrix with one row per harmonic, scale 0 to 1 (overrides all other rolloff parameters)

glottis

duration of the closed phase of a glottal cycle (silent) in relation to the open phase, % (0 = no closed phase, 100 = closed phase as long as open phase, 200 = twice as long as the open phase, etc.); numeric vector or dataframe specifying time and value (anchor format). Use this effect sparingly: it is slow, and high values may affect harmonic composition and require a high sampling rate, especially in combination with high pitch

pitchFloor, pitchCeiling

lower & upper bounds of f0

pitchSamplingRate

sampling frequency of the pitch contour only, Hz. Low values reduce processing time. Set to pitchCeiling for optimal speed or to samplingRate for optimal quality

noise

intensity of turbulent noise (0 dB = same RMS as that of the periodic (voiced) component, negative values = less intense; anchor format). In soundgen 3.0, the noise component is always calibrated relative to the filtered harmonic component. When noise is present, the harmonic and noise components are filtered separately, their RMS amplitudes are normalized after filtering, and they are then mixed. Because noise can begin before the voiced part and continue after it, the time of noise anchors MUST be in ms, not [0, 1]; this is different from all other soundgen arguments that accept the anchor format with time either in ms or [0, 1]

rolloffNoise, rolloffNoiseExp, noiseFlatSpec

linear (rolloffNoise, dB/kHz, anchor format) or exponential (rolloffNoiseExp, dB/oct, anchor format) rolloff of the excitation source for the noise component (anchor format) applied above noiseFlatSpec (Hz, scalar). More negative rolloff = less high-frequency noise

formants

a vector of formant frequencies (assuming formants are static throughout the sound); a list of formant times, frequencies, amplitudes, and bandwidths; or a character string referring to default presets for speaker "M1" (implemented: "aoieu0"). NA or NULL means no formants, only lip radiation (but a schwa is generated if vocalTract is specified). Time stamps for formants and mouth can be specified in ms relative to sylLen or on a scale of [0, 1]. See getFormantFilter for more details

formantsNoise

the same as formants, but for the aperiodic noise rather than for the periodic component. If NA (default), the noise will be filtered through the same formants as the periodic component, approximating aspiration noise [h]

formantDep

scale factor of formant amplitude (1 = no change relative to amplitudes in formants)

formantDepStoch

the amplitude of additional stochastic formants added above the highest specified formant, dB (only if temperature > 0)

formantWidth

scale factor of formant bandwidth (1 = no change)

formantCeiling

frequency to which stochastic formants are calculated to avoid losing energy in the upper part of the spectrum due to unmodeled resonances above the Nyquist frequencies, specified in multiples of Nyquist. If NULL (default), a faster theoretical correction is used (may fail for unusual sounds)

formantLocking

the approximate proportion of sound in which one of the harmonics is locked to the nearest formant, 0 = none, 1 = the entire sound (anchor format). In multi-syllable sounds, formant locking is applied separately to each syllable using the corresponding portion of the bout-level formant filter

vocalTract

the length of vocal tract, cm. Used for calculating formant dispersion (for adding extra formants) and formant transitions as the mouth opens and closes. If NULL or NA, the length is estimated based on specified formant frequencies, if any (anchor format)

mouth

mouth opening (0 to 1, 0.5 = neutral, i.e. no modification) (anchor format)

lipRad

the effect of lip radiation on source spectrum, dB/oct (the default of +6 dB/oct produces a high-frequency boost when the mouth is open)

noseRad

the effect of radiation through the nose on source spectrum, dB/oct (the alternative to lipRad when the mouth is closed)

mouthOpenThres

open the lips (switch from nose radiation to lip radiation) when the mouth is open >mouthOpenThres, 0 to 1

amDep

amplitude modulation (AM) depth, %. 0: no change; 100: AM with amplitude range equal to the dynamic range of the sound (anchor format)

amFreq

AM frequency, Hz (anchor format)

amType

"logistic" = logistic (default), "sine" = sinusoidal

amShape

ignored if amType = "sine", otherwise determines the shape of non-sinusoidal AM: 0 = ~sine, -1 = notches, +1 = clicks (anchor format)

vibratoFreq

the rate of regular pitch modulation, or vibrato, Hz (anchor format)

vibratoDep

the depth of vibrato, semitones (anchor format)

jitterDep

cycle-to-cycle random pitch variation, semitones (anchor format)

jitterLen

duration of stable periods between pitch jumps, ms. Use a low value for harsh noise, a high value for irregular vibrato or shaky voice (anchor format)

shimmerDep

random variation in amplitude between individual glottal cycles (0 to 100% of original amplitude of each cycle) (anchor format)

shimmerLen

duration of stable periods between amplitude jumps, ms. Use a low value for harsh noise, a high value for shaky voice (anchor format)

subRatio

a positive integer giving the ratio of f0 (the main fundamental) to g0 (a lower frequency): 1 = no subharmonics, 2 = period doubling regardless of pitch changes, 3 = period tripling, etc.

subDep

the depth of subharmonics relative to the main frequency component (f0), %. 0: no subharmonics; 100: g0 harmonics are as strong as the nearest f0 harmonic (anchor format)

nonlinBalance

hyperparameter for regulating the (approximate) proportion of sound with different regimes of pitch effects (none / subharmonics only / subharmonics and jitter). 0% = no nonlinear phenomena; 100% = the entire sound has jitter + subharmonics. Ignored if temperature = 0

nonlinRandomWalk

a numeric vector specifying the timing of nonlinear regimes: 0 = none, 1 = subharmonics, 2 = subharmonics + jitter + shimmer

shortestEpoch

minimum duration of each epoch with unchanging subharmonics regime or formant locking, in ms

temperature

hyperparameter for regulating the amount of stochasticity in sound generation

tempEffects

a list of scaling coefficients regulating the effect of temperature on particular parameters. To change, specify just those pars that you want to modify (1 = default, 0 = no stochastic behavior).

amplDep, pitchDep, noiseDep

random fluctuations of user-specified amplitude / pitch / noise anchors

amplDriftDep

drift of amplitude mirroring pitch drift

formDisp

dispersion of stochastic formants

formDrift

formant frequencies

glottisDep

proportion of glottal cycle with closed glottis

pitchDriftDep

amount of slow random drift of f0

pitchDriftFreq

frequency of slow random drift of f0

rolloffDriftDep

drift of rolloff mirroring pitch drift

specDep

rolloff, rolloffNoise, nonlinear effects, attack

subDriftDep

drift of subharmonic frequency and bandwidth mirroring pitch drift

sylLenDep

duration of syllables and pauses

maleFemale

hyperparameter for shifting f0 contour, formants, and vocalTract to make the speaker appear more male (-1...0) or more female (0...+1); 0 = no change

creakyBreathy

hyperparameter for an adjustment of voice quality from creaky (-1) to breathy (+1); 0 = no change

plot

if TRUE, plots a spectrogram

play

if TRUE, plays the synthesized sound using the default player on your system. If character, passed to play as the name of player to use, eg "aplay", "play", "vlc", etc. In case of errors, try setting another default player for play

saveAudio

if TRUE, saves the result as "soundgen.wav" in the working directory

invalidArgAction

what to do if an argument is invalid or outside the permitted range: 'adjust' = reset to default value, 'abort' = stop execution, 'ignore' = throw a warning and continue (may crash)

smoothing

a list of parameters passed to interpolate to control the interpolation and smoothing of contours drawn through anchors

samplingRate

sampling rate of the output (Hz)

windowLength

length of the FFT window (ms)

step

step between successive windows (ms); if provided, overrides overlap

overlap

overlap between successive windows (0–100%)

wn

wn window type accepted by winFun: character string or function

dynamicRange

dynamic range (dB). Harmonics and noise more than dynamicRange under maximum amplitude are discarded to save computational resources

...

other plotting parameters passed to spectrogram

Value

The synthesized waveform as a numeric vector normalized to [-1, 1] and sampled at the specified samplingRate. Note that the sampling rate may be increased internally in case the pitch is high, with a warning.

Parameter groups

Temporal structure

repeatBout, nSyl, sylLen, pauseLen, addSilence

Amplitude

ampl, amplGlobal, attackLen

Pitch & periodic source

pitch, pitchGlobal, rolloff, rolloffOct, rolloffKHz, rolloffExact, glottis, pitchFloor, pitchCeiling, pitchSamplingRate

Aperiodic source

noise, rolloffNoise, rolloffNoiseExp, noiseFlatSpec

Filter

formants, formantsNoise, formantDep, formantDepStoch, formantWidth, formantCeiling, formantLocking, vocalTract, lipRad, noseRad, mouth, mouthOpenThres

Amplitude & frequency modulation

amDep, amFreq, amType, amShape, vibratoFreq, vibratoDep

Nonlinear phenomena

jitterDep, jitterLen, shimmerDep, shimmerLen, subRatio, subDep, nonlinBalance, nonlinRandomWalk, shortestEpoch

Stochasticity & hyper-parameters

temperature, tempEffects, maleFemale, creakyBreathy

I/O

plot, play, saveAudio, ...

Technical

smoothing, invalidArgAction, samplingRate, windowLength, step, overlap, wn, dynamicRange

Anchor format

soundgen() and some other functions in this package accept arguments that define a time series (e.g., pitch or amplitude contour). These arguments can be provided as either numeric vectors or a list of anchors - points through which the contour is interpolated. This "anchor format" can be a dataframe or list with two elements: $time (ms or 0 to 1) and value of each anchor. Ex.: soundgen(pitch = list(time = c(0, 0.2, 1), value = c(310, 340, 280))).

See Also

getFormantFilter beat

Examples

# Detailed documentation: https://cogsci.se/soundgen/sound_generation.html

# A gallery of examples with code: https://cogsci.se/soundgen/demos.html

# A GUI for soundgen is available as a Shiny app.
# Type "soundgen_app()" to open it in your default browser

# Set "playback" to TRUE for default system player or the name of preferred
# player (eg "aplay") to play back the audio from examples
playback = FALSE # or TRUE, 'aplay', 'vlc', etc. (see ?playme)

sound = soundgen(play = playback)
# spectrogram(sound, 16000)
# playme(sound)

# Control of intonation, amplitude envelope, formants
s0 = soundgen(
  pitch = c(300, 390, 250),
  ampl = data.frame(time = c(0, 50, 300), value = c(-5, -10, 0)),
  attackLen = c(10, 50),
  formants = c(600, 900, 2200),
  play = playback
)

# Use the in-built collection of presets:
# names(presets)  # speakers
# names(presets$Chimpanzee)  # calls per speaker
s1 = eval(parse(text = presets$Chimpanzee$Scream_conflict))  # screaming chimp
# playme(s1)
s2 = eval(parse(text = presets$F1$Scream))  # screaming woman
# playme(s2, 18320)

# presets of some vowels and consonants
names(presets$M1$Formants$vowels)
soundgen(sylLen = 500, formants = 'aoieu0', play = playback)

## Not run: 
# Ultrasound - need to adjust some defaults:
 soundgen(
   sylLen = 10,  # just 10 ms
   attackLen = 1,  # should be very short for short vocalizations
   addSilence = 2,
   pitch = c(45000, 35000, 65000, 60000),  # 35-60 kHz
   rolloff = -12,
   rolloffKHz = 0,
   formants = NA,  # no formants (or set vocal tract length)
   samplingRate = 350000,  # at least ~10 times the max f0
   pitchSamplingRate = 350000,  # the same as samplingRate
   windowLength = .25,  # need very short window lengths for USV
   pitchCeiling = 90000, # max allowed pitch
   invalidArgAction = 'ignore', # override the ranges allowed by default
   temperature = 1e-4,
   plot = TRUE
 )

# soundgen plays Bach
dt = otherToHz(
  c('E4', 'D4', 'E4', 'C4', 'E4', 'B3', 'E4', 'A3', 'E4', 'G#3', 'E4',
  'A3', 'E4', 'B3', 'E4', 'C4', 'E4', 'E3', 'E4', 'F#3', 'E4', 'G#3',
  'E4', 'A3', 'E4', 'G#3', 'E4', 'A3', 'E4', 'B3', 'E4', 'C4'), 'notes')
out = numeric(0)
for (s in 1:length(dt)) {
  syl_s = soundgen(sylLen = 100, pitch = dt[s], rolloff = -15,
                   formants = c(750, 1400, 2900, 3800), noise = -45,
                   attackLen = 50, addSilence = 0, temperature = .01)
  syl_s = fade(soundgen:::matchLengths(syl_s, 0.1 * 16000), samplingRate = 16000)
  out = c(out, syl_s[1:(0.1 * 16000)])
}
spectrogram(out, 16000, yScale = 'ERB')
playme(out, 16000)

## End(Not run)

soundgen documentation built on Sept. 20, 2026, 5:07 p.m.