View source: R/shiftFormants.R
| shiftFormants | R Documentation |
Raises or lowers formants (resonance frequencies), changing the voice quality
or timbre of the sound without changing its pitch, statically or dynamically.
Note that this is only possible when the fundamental frequency f0 is lower
than the formant frequencies. For best results, freqWindow should be
no lower than f0 and no higher than formant bandwidths. Obviously, this is
impossible for many signals, so just try a few reasonable values, like ~200
Hz for speech. If freqWindow is not specified, it is set to the
median detected f0, which is slow and requires detectable pitch; if pitch
detection fails, freqWindow defaults to 400 Hz with a message.
shiftFormants(
x,
multFormants,
samplingRate = NULL,
freqWindow = NULL,
dynamicRange = 80,
windowLength = 50,
step = NULL,
overlap = 75,
wn = "gaussian",
specEnvMethod = c("cepstral", "gauss", "movavg", "peak"),
interpol = "splineFC",
normalize = c("orig", "max", "none"),
play = FALSE,
saveAudio = FALSE,
reportEvery = NULL,
cores = 1
)
x |
path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors |
multFormants |
1 = no change, >1 = raise formants (eg 1.1 = 10% up, 2 =
one octave up), <1 = lower formants. Anchor format accepted (see
|
samplingRate |
sampling rate of |
freqWindow |
the width of spectral smoothing window, Hz: a single positive number, roughly the expected spacing between harmonics. Defaults to median detected f0 |
dynamicRange |
regions under |
windowLength |
length of the analysis window, ms |
step |
step between successive windows, ms; if provided, overrides
|
overlap |
overlap between successive windows, % |
wn |
window type accepted by |
specEnvMethod |
the method of extracting a smoothed spectral envelope:
"cepstral" = Gaussian liftering of the real cepstrum (default), "gauss" =
Gaussian blur of the log spectrum, "movavg" = moving average of the log
spectrum, "peak" = moving maximum (upper envelope). See
|
interpol |
the method for interpolating scaled spectra: any method
supported by |
normalize |
"orig" = same as input (default), "max" = maximum possible peak amplitude for the given input scale, "none" = no normalization |
play |
if TRUE, plays the output audio using the default player on your
system. If a character string, it is passed to |
saveAudio |
if TRUE, saves the processed audio in a subdirectory named after the function and created in the input directory (if input is a file or folder) or in the working directory |
reportEvery |
when processing multiple inputs, report estimated time
left every |
cores |
number of cores for parallel processing |
Algorithm: phase vocoder. In the frequency domain, we separate the complex
spectrum of each STFT frame into two parts. The "receiver" is the flattened
or smoothed complex spectrum, where smoothing is achieved by obtaining a
smoothed envelope of the magnitude spectrum (the amount of smoothing is
controlled by freqWindow) and then dividing the complex spectrum by
this envelope. To avoid amplifying noise, bins more than dynamicRange
dB below the frame peak are divided by a floored envelope. This division
basically removes the formants from the signal. The second component,
"donor", is a scaled and interpolated version of the same smoothed magnitude
envelope as above - these are the formants shifted up or down. Warping can be
easily implemented instead of simple scaling if nonlinear spectral
transformations are required. We then multiply the "receiver" and "donor"
spectrograms and reconstruct the audio with iSTFT.
The processed waveform as a numeric vector with the original sampling rate (a list if there are multiple inputs).
shiftPitch transplantFormants
getSpecEnv
data('speechEx', package = 'soundgen')
# playme(speechEx)
# spectrogram(speechEx)
# Lower formants by 4 semitones or ~20% = 2 ^ (-4 / 12)
speech1 = shiftFormants(speechEx, multFormants = 2 ^ (-4 / 12), freqWindow = 150)
# playme(speech1, speechEx@samp.rate)
# spectrogram(speech1, speechEx@samp.rate)
orig = meanSpectrum(speechEx, plot = FALSE)
shifted = meanSpectrum(speech1, speechEx@samp.rate, plot = FALSE)
plot(ampl ~ freq, orig, log = 'y', type = 'l')
lines(ampl ~ freq, shifted, col = 'blue')
# dynamic change: raise formants at the beginning, lower at the end
speech2 = shiftFormants(speechEx, multFormants = c(1.3, .7), freqWindow = 150)
# playme(speech2, speechEx@samp.rate)
# spectrogram(speech2, speechEx@samp.rate)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.