shiftPitch: Shift pitch

View source: R/shiftPitch.R

shiftPitchR Documentation

Shift pitch

Description

Raises or lowers pitch with or without also shifting the formants (resonance frequencies) and performing a time-stretch. The three operations (pitch shift, formant shift, and time stretch) are independent and can be performed in any combination, statically or dynamically. shiftPitch can also be used to shift formants without changing pitch or duration, but the dedicated shiftFormants is faster for that task. Likewise, use the much faster timeStretch for slowing down or speeding up a recording without preserving either pitch or formants. Tip: increase overlap to >90% for best quality.

Usage

shiftPitch(
  x,
  samplingRate = NULL,
  multPitch = 1,
  multFormants = multPitch,
  timeStretch = 1,
  freqWindow = NULL,
  dynamicRange = 80,
  windowLength = 40,
  step = NULL,
  overlap = 90,
  wn = "hanning",
  interpol = "splineFC",
  propagation = c("time", "adaptive"),
  specEnvMethod = c("cepstral", "gauss", "movavg", "peak"),
  preserveEnv = FALSE,
  transplantEnv_pars = list(windowLength = 10),
  normalize = c("orig", "max", "none"),
  play = FALSE,
  saveAudio = FALSE,
  reportEvery = NULL,
  cores = 1
)

Arguments

x

path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors

samplingRate

sampling rate of x (only needed if x is a numeric vector)

multPitch

1 = no change, >1 = raise pitch (eg 1.1 = 10% up, 2 = one octave up), <1 = lower pitch. Anchor format accepted for multPitch / multFormants / timeStretch (see soundgen)

multFormants

1 = no change, >1 = raise formants (eg 1.1 = 10% up, 2 = one octave up), <1 = lower formants. The default behavior is for formants to follow pitch with multFormants = multPitch

timeStretch

1 = no change, >1 = longer, <1 = shorter

freqWindow

the width of spectral smoothing window, Hz - see shiftFormants for discussion and examples. Defaults to detected f0 prior to pitch shifting times mean(multPitch). Note that formant shifting is applied AFTER pitch shifting, so freqWindow should be multiplied by multPitch compared to what you would use on the original audio

dynamicRange

dynamic range (dB) that silences spectrogram regions below -dyamicRange dB and also control the tol parameter in istft_timevar

windowLength

length of the analysis window, ms

step

step between successive windows, ms; if provided, overrides overlap; because digital audio is sampled at discrete time intervals of 1/samplingRate, the actual step and thus the time stamps of STFT frames may be slightly different - e.g., 24.98866 instead of 25.0 ms

overlap

overlap between successive windows, %

wn

window type accepted by winFun: character string or function

interpol

the method for interpolating scaled spectra and anchors: any method supported by interpolate, defaults to "splineFC"

propagation

the method for propagating phase: "time" (default) = single-pass horizontal propagation, several times faster; "adaptive" = a modified "vocoder done right" (Prusa & Holighaus 2017), high-quality but relatively slow

specEnvMethod

the method of extracting a smoothed spectral envelope: "cepstral" = Gaussian liftering of the real cepstrum (default), "gauss" = Gaussian blur of the log spectrum, "movavg" = moving average of the log spectrum, "peak" = moving maximum (upper envelope). See getSpecEnv for details

preserveEnv

if TRUE, transplants the amplitude envelope from the original to the modified sound with transplantEnv. Mostly this makes sense if there is no time stretching

transplantEnv_pars

a list of parameters passed on to transplantEnv if preserveEnv = TRUE

normalize

"orig" = same as input (default), "max" = maximum possible peak amplitude for the given input scale, "none" = no normalization

play

if TRUE, plays the output audio using the default player on your system. If a character string, it is passed to playme as the name of the player to use (e.g. 'aplay', 'play', 'vlc'). In case of errors, try setting another default player for playme

saveAudio

if TRUE, saves the processed audio in a subdirectory named after the function and created in the input directory (if input is a file or folder) or in the working directory

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

cores

number of cores for parallel processing

Details

Algorithm: phase vocoder. Pitch shifting is accomplished by performing a time stretch (at present, with horizontal or adaptive phase propagation) followed by resampling. This shifts both pitch and formants; to preserve the original formant frequencies or modify them independently of pitch, the recipient's spectral envelope is flattened and the donor envelope is imposed on the same spectrogram before iSTFT. See Prusa 2017 "Phase vocoder done right", Royer 2019 "Pitch-shifting algorithm design and applications in music".

Value

The processed waveform as a numeric vector (a list if there are multiple inputs).

See Also

shiftFormants transplantFormants

Examples

data(speechEx, package = 'soundgen')
samplingRate = speechEx@samp.rate
# playme(speechEx)
spectrogram(speechEx, yScale = 'ERB')

# raise pitch, lower formants
s1 = shiftPitch(speechEx, freqWindow = 200, multPitch = 1.2, multFormants = .85)
# spectrogram(s1, samplingRate, yScale = 'ERB')
# playme(s1, samplingRate)

## Not run: 
# Tips for best quality: high overlap (slow), pad with silence
# and fade a bit before processing; try adaptive phase propagation (slow)
s = c(
  rep(0, 1000),
  fade(speechEx@left, samplingRate, fadeIn = 50, fadeOut = 50),
  rep(0, 1000)
)
s1a = shiftPitch(s, samplingRate, freqWindow = 200, multPitch = 1.2,
  multFormants = .85, propagation = 'adaptive', overlap = 95)
spectrogram(s1a, samplingRate, yScale = 'ERB')
playme(s1a, samplingRate)
# cf:
playme(s1, samplingRate)

## Dynamic manipulations
# Add a chevron-shaped contour to both pitch and formants
s2 = shiftPitch(speechEx, multPitch = c(1.1, 1.3, .8))
playme(s2, samplingRate)
spectrogram(s2, samplingRate, yScale = 'ERB')

# Time-stretch only the middle
s3 = shiftPitch(speechEx, overlap = 95, timeStretch = list(
  time = c(0, .25, .31, .5, .55, 1),
  value = c(1, 1, 3, 3, 1, 1))
)
playme(s3, samplingRate)

# Raise pitch and formants by 3 semitones, shorten
s4 = shiftPitch(speechEx, multPitch = 2 ^ (3 / 12), timeStretch = 0.75)
playme(s4, samplingRate)
spectrogram(s4, samplingRate, yScale = 'ERB')

# Just speed up
shiftPitch(speechEx, multPitch = 1, timeStretch = 0.75, play = TRUE)

# Raise pitch, preserve formants
s5 = shiftPitch(speechEx, multPitch = 1.2, multFormants = 1, freqWindow = 150)
playme(s5, samplingRate)
spectrogram(s5, samplingRate, yScale = 'ERB')

# Only modify voiced frames, preserving consonants / breathing etc
s1 = soundgen(nSyl = 2, sylLen = 300, pauseLen = 500, pitch = c(250, 200))
s2 = soundgen(sylLen = 150, noise = 0, pitch = NA, formants = list(f1 = 5000))
s = addVectors(s1, s2, insertionPoint = 16000 * .5)
s = s + rnorm(length(s), 0, .01)
spectrogram(s, 16000, yScale = 'ERB')
playme(s, 16000)  # we want to ignore the central /tch/ when shifting f0
# run analyze to get f0 contours, correct manually with pitch_app if needed
an = analyze(s, 16000, windowLength = 25, step = 5, plot = TRUE, yScale = 'ERB')
pitch = an$detailed$pitch
multPitch = ifelse(is.na(pitch), 1, 1.25)
s_shifted = fade(shiftPitch(s, 16000, multPitch = multPitch, wn = 'hanning', overlap = 90))
spectrogram(s_shifted, 16000, yScale = 'ERB')
playme(s_shifted, 16000)  # the /tch/ in the middle is unaffected by pitch-shifting

## End(Not run)

soundgen documentation built on Sept. 20, 2026, 5:07 p.m.