shiftFormants: Shift formants

View source: R/shiftFormants.R

shiftFormantsR Documentation

Shift formants

Description

Raises or lowers formants (resonance frequencies), changing the voice quality or timbre of the sound without changing its pitch, statically or dynamically. Note that this is only possible when the fundamental frequency f0 is lower than the formant frequencies. For best results, freqWindow should be no lower than f0 and no higher than formant bandwidths. Obviously, this is impossible for many signals, so just try a few reasonable values, like ~200 Hz for speech. If freqWindow is not specified, it is set to the median detected f0, which is slow and requires detectable pitch; if pitch detection fails, freqWindow defaults to 400 Hz with a message.

Usage

shiftFormants(
  x,
  multFormants,
  samplingRate = NULL,
  freqWindow = NULL,
  dynamicRange = 80,
  windowLength = 50,
  step = NULL,
  overlap = 75,
  wn = "gaussian",
  specEnvMethod = c("cepstral", "gauss", "movavg", "peak"),
  interpol = "splineFC",
  normalize = c("orig", "max", "none"),
  play = FALSE,
  saveAudio = FALSE,
  reportEvery = NULL,
  cores = 1
)

Arguments

x

path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors

multFormants

1 = no change, >1 = raise formants (eg 1.1 = 10% up, 2 = one octave up), <1 = lower formants. Anchor format accepted (see soundgen)

samplingRate

sampling rate of x (only needed if x is a numeric vector)

freqWindow

the width of spectral smoothing window, Hz: a single positive number, roughly the expected spacing between harmonics. Defaults to median detected f0

dynamicRange

regions under -dynamicRange dB are treated as silent

windowLength

length of the analysis window, ms

step

step between successive windows, ms; if provided, overrides overlap; because digital audio is sampled at discrete time intervals of 1/samplingRate, the actual step and thus the time stamps of STFT frames may be slightly different - e.g., 24.98866 instead of 25.0 ms

overlap

overlap between successive windows, %

wn

window type accepted by winFun: character string or function

specEnvMethod

the method of extracting a smoothed spectral envelope: "cepstral" = Gaussian liftering of the real cepstrum (default), "gauss" = Gaussian blur of the log spectrum, "movavg" = moving average of the log spectrum, "peak" = moving maximum (upper envelope). See getSpecEnv for details

interpol

the method for interpolating scaled spectra: any method supported by interpolate, defaults to "splineFC"

normalize

"orig" = same as input (default), "max" = maximum possible peak amplitude for the given input scale, "none" = no normalization

play

if TRUE, plays the output audio using the default player on your system. If a character string, it is passed to playme as the name of the player to use (e.g. 'aplay', 'play', 'vlc'). In case of errors, try setting another default player for playme

saveAudio

if TRUE, saves the processed audio in a subdirectory named after the function and created in the input directory (if input is a file or folder) or in the working directory

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

cores

number of cores for parallel processing

Details

Algorithm: phase vocoder. In the frequency domain, we separate the complex spectrum of each STFT frame into two parts. The "receiver" is the flattened or smoothed complex spectrum, where smoothing is achieved by obtaining a smoothed envelope of the magnitude spectrum (the amount of smoothing is controlled by freqWindow) and then dividing the complex spectrum by this envelope. To avoid amplifying noise, bins more than dynamicRange dB below the frame peak are divided by a floored envelope. This division basically removes the formants from the signal. The second component, "donor", is a scaled and interpolated version of the same smoothed magnitude envelope as above - these are the formants shifted up or down. Warping can be easily implemented instead of simple scaling if nonlinear spectral transformations are required. We then multiply the "receiver" and "donor" spectrograms and reconstruct the audio with iSTFT.

Value

The processed waveform as a numeric vector with the original sampling rate (a list if there are multiple inputs).

See Also

shiftPitch transplantFormants getSpecEnv

Examples

data('speechEx', package = 'soundgen')
# playme(speechEx)
# spectrogram(speechEx)

# Lower formants by 4 semitones or ~20% = 2 ^ (-4 / 12)
speech1 = shiftFormants(speechEx, multFormants = 2 ^ (-4 / 12), freqWindow = 150)
# playme(speech1, speechEx@samp.rate)
# spectrogram(speech1, speechEx@samp.rate)

orig = meanSpectrum(speechEx, plot = FALSE)
shifted = meanSpectrum(speech1, speechEx@samp.rate, plot = FALSE)
plot(ampl ~ freq, orig, log = 'y', type = 'l')
lines(ampl ~ freq, shifted, col = 'blue')

# dynamic change: raise formants at the beginning, lower at the end
speech2 = shiftFormants(speechEx, multFormants = c(1.3, .7), freqWindow = 150)
# playme(speech2, speechEx@samp.rate)
# spectrogram(speech2, speechEx@samp.rate)

soundgen documentation built on Sept. 20, 2026, 5:07 p.m.