View source: R/modulationSpectrum.R
| modulationSpectrum | R Documentation |
Produces a modulation spectrum of waveform(s) or audio file(s). It begins
with a spectrogram-like time-frequency representation and analyzes the
modulation of the envelope in each frequency band. If specFun =
'audSpec', the sound is passed through a bank of bandpass filters with
audSpectrogram. If specFun = 'STFT', we begin with an
ordinary spectrogram produced with a Short-Time Fourier Transform. If
msType = '2D', the modulation spectrum is a 2D Fourier transform of
the spectrogram-like representation, with temporal modulation along the X
axis and spectral modulation along the Y axis. A good visual analogy is
decomposing the spectrogram into a sum of ripples of various frequencies and
directions. If msType = '1D', the modulation spectrum is a matrix
containing 1D Fourier transforms of each frequency band in the spectrogram,
so the result again has modulation frequencies along the X axis, but the Y
axis now shows the frequency of each analyzed band.
modulationSpectrum(
x,
samplingRate = NULL,
scale = NULL,
from = NULL,
to = NULL,
msType = c("2D", "1D"),
specFun = "STFT",
specFun_pars = list(),
amRes = 5,
maxDur = 5,
specMethod = c("meanSpectrum", "spectrum"),
logSpec = FALSE,
logMPS = FALSE,
power = 1,
normalize = TRUE,
roughRange = NULL,
roughMean = 100,
roughSD = 8,
roughMinFreq = 1,
amRange = c(10, 200),
fluctRange = c(0.25, 30),
fluctMean = 4,
fluctSD = 12,
returnMS = TRUE,
returnComplex = FALSE,
summaryFun = c("mean", "sd"),
output = "all",
averageMS = FALSE,
reportEvery = NULL,
cores = 1,
plot = TRUE,
savePlots = FALSE,
embed = FALSE,
logWarpX = NULL,
logWarpY = NULL,
quantiles = c(0.5, 0.8, 0.9),
kernelSize = 5,
kernelSD = 0.5,
colorTheme = "bw",
col = NULL,
main = NULL,
xlab = "Hz",
ylab = NULL,
xlim = NULL,
ylim = NULL,
width = 900,
height = 500,
units = "px",
res = NA,
...
)
x |
path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors |
samplingRate |
sampling rate of |
scale |
maximum possible amplitude of input, used to normalize the input
vector (only needed if |
from, to |
if NULL (default), analyzes the whole sound, otherwise from...to (s) |
msType |
'2D' = two-dimensional Fourier transform of a spectrogram; '1D' = separately calculated spectrum of each frequency band |
specFun |
'STFT' or 'stft' = ordinary |
specFun_pars |
a list of parameters passed to |
amRes |
approximate frequency resolution of amplitude modulation, Hz;
equivalently, the number of independent MS / roughness / fluctuation values
extracted per second of audio. Thus, larger |
maxDur |
sounds longer than |
specMethod |
the function to call when calculating the spectrum of each
frequency band, only used when |
logSpec |
if |
logMPS |
if |
power |
numeric exponent applied to the modulation spectrum, e.g.
|
normalize |
if |
roughRange |
the range of temporal modulation frequencies that
constitute the "roughness" zone, Hz. If |
roughMean, roughSD |
the mean (Hz) and standard deviation (semitones) of
a lognormal distribution used to weight roughness estimates. Weighted
roughness requires both |
roughMinFreq |
frequencies below |
amRange |
the range of temporal modulation frequencies in which we look for systematic amplitude modulation, Hz |
fluctRange, fluctMean, fluctSD |
same as |
returnMS |
if |
returnComplex |
if |
summaryFun |
summary functions used to summarize per-fragment roughness,
|
output |
what to return, options: 'original', 'original_list', 'modulation_spectrogram', 'processed', 'complex', 'roughness', 'roughness_spectrogram', 'roughness_list', 'fluctuation', 'fluctuation_list', 'fluctuation_spectrogram', 'amMsFreq', 'amMsPurity', 'ampl', 'all' (see the Return section) |
averageMS |
if |
reportEvery |
when processing multiple inputs, report estimated time
left every |
cores |
number of cores for parallel processing |
plot |
if |
savePlots |
if TRUE, creates a subdirectory in the input directory (if input is a file or folder) or in the working directory (if input is a vector etc), named after the function (eg "spectrogram/"). All plots and audio files (if any) are saved in this new directory. If there are multiple inputs, an html notebook is also created for easy viewing and listening |
embed |
if TRUE and savePlots is set and there are multiple inputs, all saved images and audio (if any) are embedded in the exported html notebook for easy sharing; if FALSE, the html file links to separate images and audio files (but separate files are still saved). NB: for this to work, package "base64enc" must be installed |
logWarpX, logWarpY |
numeric vector of length 2: |
quantiles |
numeric vector of cumulative contours in (0, 1). For
example, |
kernelSize |
the size of the Gaussian kernel used for smoothing the modulation spectrum prior to plotting; 1 means no smoothing |
kernelSD |
the SD of the Gaussian kernel used for smoothing, relative to its size |
colorTheme |
plot color theme |
col |
custom color palette; if supplied, overrides |
main, xlab, ylab, xlim, ylim |
graphical parameters |
width, height, units, res |
parameters passed to
|
... |
other graphical parameters passed on to
|
To calculate roughness and fluctuation depth, the symmetric modulation
spectrum is folded by averaging its ±AM halves. Roughness is calculated as
the proportion of the modulation spectrum within roughRange of
temporal modulation frequencies or as a weighted version thereof. Fluctuation
is calculated in the same manner as roughness, but in a lower frequency range
of ~4 Hz (0.25 to 30 Hz). The frequency of amplitude modulation
(amMsFreq, Hz) is calculated as the highest peak of the folded AM
function within amRange, and its purity (amMsPurity, dB) as the
ratio of this peak to the average of the folded AM function around the peak
within amRange. For relatively short and steady sounds, set amRes =
NULL and analyze the entire sound. For longer sounds and when roughness or
AM vary over time, set amRes to get multiple measurements over time,
but inspect the MS visually to make sure the relevant modulation frequencies
are resolved. For multiple inputs, such as a list of waveforms or a path to a
folder with audio files, the ensemble of modulation spectra can be
interpolated to the same spectral and temporal resolution and averaged if
averageMS = TRUE.
The full STFT or auditory spectrogram is computed once and then sliced into
temporal chunks. Roughness, fluctuation strength, amMsFreq, and
amMsPurity are calculated from the per-fragment modulation spectra
before Gaussian smoothing and before any log-warping used for plotting.
Log-warps specified via logWarpX and logWarpY affect the plot
only; the returned matrices are not log-warped.
A list with three top-level elements: $summary,
$detailed, and $aggregated.
$summary is a dataframe summarizing roughness, amMsFreq,
and amMsPurity for each input (one row per file), or NULL
if summaryFun = NULL.
$detailed contains detailed statistics per sound. If multiple
sounds are analyzed, it is a named list of per-sound lists. If a single
sound is analyzed, it is simplified to a single list. Components are
selected via the output argument and may include:
averaged modulation spectrum across fragments,
after logMPS, power, and normalize if requested, but
before Gaussian smoothing and plotting transforms. Colnames are temporal
modulation frequencies (Hz). Rownames are spectral modulation frequencies
if msType = '2D' and frequencies of filters or spectrogram bands if
msType = '1D'.
list of per-fragment modulation spectra.
a spectrogram-like representation
showing how the frequency-averaged modulation spectrum changes over time;
NA if there is only one fragment.
modulation spectrum after Gaussian smoothing. This matrix is not log-warped; log-warping is applied only during plotting.
complex modulation spectrum, returned only if
returnComplex = TRUE, returnMS = TRUE, and
msType = '2D'.
proportion of the modulation spectrum within
roughRange, or a weighted version thereof if roughMean and
roughSD are supplied, in percent. This is a vector if the sound is
analyzed in multiple fragments, otherwise a single number.
a list containing frequencies, band-wise roughness values, and band amplitudes for each fragment.
a spectrogram-like matrix showing roughness per frequency band and per fragment.
proportion of the modulation spectrum within
fluctRange, or a weighted version thereof if fluctMean and
fluctSD are supplied, in percent. This is a vector if the sound is
analyzed in multiple fragments, otherwise a single number.
a list containing frequencies, band-wise fluctuation values, and band amplitudes for each fragment.
a spectrogram-like matrix showing fluctuation strength per frequency band and per fragment.
frequency of the highest AM peak within
amRange. Like roughness, this can be a single number or a
vector, depending on whether the sound is analyzed as a whole or in chunks.
ratio of the AM peak at amMsFreq to the
mean amplitude of the other folded AM values within amRange, in dB.
Root Mean Square amplitude of each analyzed fragment, divided by the internal audio scaling factor if present.
$aggregated is a list containing the original,
processed, and (if requested) complex modulation spectra
averaged across all successfully analyzed sounds. Only present if
averageMS = TRUE and multiple sounds are analyzed.
Singh, N. C., & Theunissen, F. E. (2003). Modulation spectra of natural sounds and ethological theories of auditory processing. The Journal of the Acoustical Society of America, 114(6), 3394-3411.
Anikin, A. (2025) Acoustic estimation of voice roughness. Attention, Perception, & Psychophysics 87: 1771–1787.
plotMS spectrogram
audSpectrogram analyze
s = soundgen(pitch = 440, amFreq = 100, amDep = 50)
ms = modulationSpectrum(s, samplingRate = 16000, amRes = NULL)
ms$detailed[c('roughness', 'amMsFreq', 'amMsPurity')] # a single value for each
ms1 = modulationSpectrum(s, samplingRate = 16000, amRes = 5)
ms1$detailed[c('roughness', 'amMsFreq', 'amMsPurity')]
# measured over time (low values of amRes mean more precision, so we analyze
# longer segments and get fewer values per sound)
# Embellish
ms = modulationSpectrum(s, samplingRate = 16000, logMPS = TRUE,
xlab = 'Temporal modulation, Hz', ylab = 'Spectral modulation, 1/kHz',
colorTheme = 'matlab', main = 'Modulation spectrum', lty = 3)
# Plot a modulation spectrogram (the peak at 100 Hz shows AM)
spectrogram(s, 16000, specManual = ms$detailed$modulation_spectrogram,
ylab = 'Modulation frequency, kHz', main = 'Modulation spectrogram')
# Plot a roughness spectrogram
spectrogram(s, 16000, specManual = ms$detailed$roughness_spectrogram,
yScale = 'ERB', main = 'Roughness spectrogram')
# 1D instead of 2D
modulationSpectrum(s, 16000, msType = '1D', quantiles = NULL,
col = soundgen:::jet.col(50), logWarpX = c(10, 2))
## Not run:
# A long sound with varying AM and a bit of chaos at the end
s_long = soundgen(sylLen = 3500, pitch = c(250, 320, 280),
amFreq = c(30, 55), amDep = c(20, 60, 40),
jitterDep = c(0, 0, 2), plot = TRUE, yScale = 'ERB')
playme(s_long)
ms = modulationSpectrum(s_long, 16000)
# plot AM over time
plot(names(ms$detailed$amMsFreq), ms$detailed$amMsFreq,
cex = ms$detailed$amMsPurity/5,
xlab = 'Time, ms', ylab = 'AM frequency, Hz')
# plot roughness over time
spectrogram(s_long, 16000, ylim = c(0, 4),
extraContour = list(x = ms$detailed$roughness / max(ms$detailed$roughness) * 4000, col = 'blue'))
# As with spectrograms, there is a tradeoff in time-frequency resolution
s = soundgen(pitch = 500, amFreq = 50, amDep = 100, sylLen = 500,
samplingRate = 44100, plot = TRUE)
# playme(s, samplingRate = 44100)
ms = modulationSpectrum(s, samplingRate = 44100, amRes = NULL,
specFun_pars = list(windowLength = 50, step = 50)) # poor temporal resolution
ms = modulationSpectrum(s, samplingRate = 44100, amRes = NULL,
specFun_pars = list(windowLength = 5, step = 1)) # poor frequency resolution
ms = modulationSpectrum(s, samplingRate = 44100, amRes = NULL,
specFun_pars = list(windowLength = 15, step = 3)) # a reasonable compromise
# Start with an auditory spectrogram instead of STFT
modulationSpectrum(s, 44100, specFun = 'audSpec', xlim = c(-100, 100))
modulationSpectrum(s, 44100, specFun = 'audSpec',
logWarpX = c(10, 2), xlim = c(-500, 500),
specFun_pars = list(nFilters = 32, filterType = 'gammatone', bandwidth = NULL))
# customize the plot
ms = modulationSpectrum(s, samplingRate = 44100,
amRes = NULL,
kernelSize = 17, # more smoothing
xlim = c(-70, 70), ylim = c(0, 4), # zoom in on the central region
quantiles = c(.25, .5, .8), # customize contour lines
col = rev(rainbow(100)), # alternative palette
logWarpX = c(10, 2), # pseudo-log transform
power = 2) # ^2
# Note the peaks at FM = 2/kHz (from "pitch = 500") and AM = 50 Hz (from
# "amFreq = 50")
# Input can be path to folder with audio files. Each file is processed
# separately, and the output can contain an MS per file...
ms1 = modulationSpectrum('~/Downloads/temp', kernelSize = 11,
plot = FALSE, averageMS = FALSE)
ms1$summary
names(ms1$detailed) # separate MS and other descriptives per file
# ...or a single MS can be calculated by averaging across all files:
ms2 = modulationSpectrum('~/Downloads/temp', kernelSize = 11,
plot = FALSE, averageMS = TRUE)
plotMS(ms2$aggregated$original)
# A sound with ~3 syllables per second and only downsweeps in F0 contour
s = soundgen(nSyl = 8, sylLen = 200, pauseLen = 100, pitch = c(300, 200))
# playme(s)
ms = modulationSpectrum(s, samplingRate = 16000, maxDur = .5,
xlim = c(-25, 25), colorTheme = 'seewave',
power = 2)
# note the asymmetry b/c of downsweeps
# "power = 2" returns squared modulation spectrum - note that this affects
# the roughness measure!
ms$detailed$roughness
# compare:
modulationSpectrum(s, samplingRate = 16000, maxDur = .5,
xlim = c(-25, 25), colorTheme = 'seewave',
power = 1)$detailed$roughness # much higher roughness
# Plotting with or without log-warping the modulation spectrum:
ms = modulationSpectrum(soundgen(), samplingRate = 16000, plot = TRUE)
ms = modulationSpectrum(soundgen(), samplingRate = 16000,
logWarpX = c(2, 2), plot = TRUE)
# logWarp and kernelSize have no effect on roughness
# because it is calculated before these transforms:
modulationSpectrum(s, samplingRate = 16000, logWarpX = c(1, 10))$detailed$roughness
modulationSpectrum(s, samplingRate = 16000, logWarpX = NA)$detailed$roughness
modulationSpectrum(s, samplingRate = 16000, kernelSize = 17)$detailed$roughness
# Log-transform the spectrogram prior to 2D FFT (affects roughness):
modulationSpectrum(s, samplingRate = 16000, logSpec = FALSE)$detailed$roughness
modulationSpectrum(s, samplingRate = 16000, logSpec = TRUE)$detailed$roughness
# Use a lognormal weighting function to calculate roughness
# (instead of just % in roughRange)
modulationSpectrum(s, 16000, roughRange = NULL,
roughMean = 75, roughSD = 3)$detailed$roughness
modulationSpectrum(s, 16000, roughRange = NULL,
roughMean = 100, roughSD = 12)$detailed$roughness
# truncate weights outside roughRange
modulationSpectrum(s, 16000, roughRange = c(30, 150),
roughMean = 100, roughSD = 1000)$detailed$roughness # very large SD
modulationSpectrum(s, 16000, roughRange = c(30, 150),
roughMean = NULL)$detailed$roughness # same as above b/c SD --> Inf
# Complex modulation spectrum with phase preserved
ms = modulationSpectrum(soundgen(), samplingRate = 16000,
returnComplex = TRUE)
plotMS(log(abs(ms$detailed$complex + 1e-8))) # note the symmetry
# compare:
plotMS(ms$detailed$original)
## End(Not run)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.