filterSoundByMS: Filter sound by modulation spectrum

View source: R/invertModulSpec.R

filterSoundByMSR Documentation

Filter sound by modulation spectrum

Description

Manipulates the modulation spectrum (MS) of a sound to attenuate or amplify certain frequencies of amplitude modulation (AM) and frequency modulation (FM). Algorithm: produces a modulation spectrum with modulationSpectrum, modifies it with filterMS, converts the modified MS back to a spectrogram, and finally inverts the spectrogram with invertSpectrogram, thus producing a sound with (approximately) the desired characteristics of the MS. Note that the last step of inverting the spectrogram introduces some noise, so the resulting MS is not precisely the same as the intermediate filtered version. In practice this means that some residual energy will still be present in the filtered-out frequency range (see examples).

Usage

filterSoundByMS(
  x,
  samplingRate = NULL,
  from = NULL,
  to = NULL,
  logSpec = TRUE,
  windowLength = 25,
  step = NULL,
  overlap = 80,
  wn = "hanning",
  amCond = NULL,
  fmCond = NULL,
  jointCond = NULL,
  action = c("stop", "pass"),
  dynamicRange = 80,
  dB = Inf,
  smooth = NULL,
  protect = c("level", "meanSpectrum"),
  warnAsymmetric = TRUE,
  initialPhase = c("orig", "spsi", "random", "zero"),
  nIter = 50,
  reportEvery = NULL,
  cores = 1,
  play = FALSE,
  saveAudio = FALSE,
  plot = TRUE,
  savePlots = FALSE,
  embed = FALSE,
  width = 900,
  height = 500,
  units = "px",
  res = NA
)

filterMS(
  ms,
  amCond = NULL,
  fmCond = NULL,
  jointCond = NULL,
  action = c("stop", "pass"),
  dB = Inf,
  smooth = NULL,
  protect = c("level", "meanSpectrum"),
  warnAsymmetric = TRUE,
  plot = TRUE
)

Arguments

x

path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors

samplingRate

sampling rate of x (only needed if x is a numeric vector)

from, to

if NULL (default), analyzes the whole sound, otherwise from...to (s)

logSpec

if TRUE, the spectrogram-like representation is log-transformed prior to estimating the modulation spectrum. This applies to both 2D and 1D modulation spectra. A small constant is added to avoid non-finite values

windowLength, step, overlap, wn

STFT parameters used to create the original spectrogram; make sure zp = 0

amCond, fmCond

condition on amplitude and frequency modulation: character string, expression, or function (see examples)

jointCond

character string, expression, or function with a joint condition on am and fm

action

'stop' or 'pass'

dynamicRange

the minimum possible value of the magnitude spectrogram (dB, must be positive or NA). 80 (default) = 80 dB below the global maximum, NA or Inf = no floor (allow zero values). Lowering the dynamic range (e.g., to 60 dB) attenuates ripple effects caused by filtering, but adds some uniform residual noise at the floor level

dB

a positive number giving the strength of effect in dB (defaults to Inf - complete removal of selected frequencies)

smooth

smoothing of the MS filter mask. NULL = default smoothing (~3 bins); NA = no smoothing; numeric vector of length 2 = smoothing bandwidth for FM and AM in native units (FM: cycles/kHz, AM: Hz). A single number is not allowed because FM and AM have different units

protect

which zero-modulation components to protect from filtering: 'level' = the (0, 0) cell (overall mean level / gain); 'meanSpectrum' = the (am = 0) column (time-averaged spectrum); 'envelope' = the (fm = 0) row (broadband temporal envelope, including any AM). TRUE protects all three; FALSE, NULL or NA protects none. NB: protecting 'envelope' re-injects temporal modulation when filtering AM; protecting 'meanSpectrum' preserves static spectral modulation when filtering FM

warnAsymmetric

if TRUE, warns if the MS filter mask is asymmetric with respect to zero AM/FM. Asymmetric filters can produce complex-valued inverse transforms and phase-related artifacts

initialPhase

initial phase estimate: "orig" (default) = phase of the original sound; "spsi" = single-pass spectrogram inversion (Beauregard et al., 2015); "random" = uniformly distributed noise; "zero" = all phases zero. If nIter = 0 and initialPhase = "orig", the original phase is used directly

nIter

the number of iterations of the GL algorithm (Griffin & Lim, 1984), 0 = don't run

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

cores

number of cores for parallel processing

play

if TRUE, plays back the reconstructed audio

saveAudio

if TRUE, saves the processed audio in a subdirectory named after the function and created in the input directory (if input is a file or folder) or in the working directory

plot

if TRUE, produces a triple plot: original MS, filtered MS, and the MS of the output sound

savePlots

if TRUE, creates a subdirectory in the input directory (if input is a file or folder) or in the working directory (if input is a vector etc), named after the function (eg "spectrogram/"). All plots and audio files (if any) are saved in this new directory. If there are multiple inputs, an html notebook is also created for easy viewing and listening

embed

if TRUE and savePlots is set and there are multiple inputs, all saved images and audio (if any) are embedded in the exported html notebook for easy sharing; if FALSE, the html file links to separate images and audio files (but separate files are still saved). NB: for this to work, package "base64enc" must be installed

width, height, units, res

parameters passed to png if the plot is saved

ms

a modulation spectrum: matrix of real or complex values, AM in columns (Hz), FM in rows (cycles/kHz)

Details

By default, the filter mask is smoothed to avoid abrupt cutoffs. Smoothing is specified in physical units: FM in cycles/kHz and AM in Hz. Set smooth = NA to disable smoothing. If logSpec = TRUE, filtering is performed on the modulation spectrum of the log-magnitude spectrogram, which is often more robust and more physiologically plausible than filtering linear magnitudes. If initialPhase = 'orig', the phase of the original sound is used. With nIter = 0, the original phase is used directly; with nIter > 0, the original phase is used to initialize the Griffin-Lim iterations.

Value

Filtered audio as a numeric vector normalized to [-1, 1] with the same sampling rate as input.

Choosing filtering settings

MS filtering operates globally on the spectrogram's envelope, so any transient whose modulation content overlaps the filtered band — onsets/offsets, plosives, amplitude steps — is partially removed together with the target, producing a ripple that decays inward from the transient. Smoothing (blurring the filter mask before applying it) shortens the ripple's tails but cannot eliminate it completely. Two parameters control the trade-off between successful filtering and introducing spectro-temporal artifacts: (1) dynamicRange: lower values shrink the silence-to-sound step in the log domain and lift the linear target above the zero-clipping knee, suppressing edge ripple at the cost of a uniform residual noise floor, most visible in silent regions, and (2) dB: finite attenuation, e.g. 20–40 dB, removes proportionally less of the transients' in-band content than dB = Inf. Use dB = Inf with a large dynamicRange for steady, sustained sounds and finite dB with dynamicRange = 60–80 for sounds with transients, such as speech. Use logSpec = TRUE for manipulating the temporal envelope / AM in a perceptually meaningful manner. logSpec = FALSE weights components by absolute magnitude and avoids amplifying quiet regions, but shows stronger edge artifacts. Choose protect = c('level', 'meanSpectrum') (the default setting) when filtering AM. Do not protect 'envelope', which contains the very AM being removed. Set protect = c('level', 'envelope') when filtering FM, so as to keep the natural amplitude contour. Drop 'meanSpectrum' only if you also want to flatten the static spectral shape. Whenever the condition includes 0 on an axis (e.g. removing slow AM with abs(am) < 3), keep the corresponding protection to avoid a collapsed, "AC-coupled" reconstruction. For mild filtering, initialPhase = 'orig' with nIter = 0 is fast and preserves the original phase; increase nIter for stronger filtering.

See Also

invertSpectrogram

Examples

# Create a sound to be filtered
s = soundgen(sylLen = 500, samplingRate = 16000,
    amFreq = 25, amDep = 50,
    addSilence = 50, plot = TRUE)
# playme(s, 16000)

# Filter to remove the rapid AM at 25 Hz
s_filt = fade(filterSoundByMS(s, samplingRate = 16000,
  amCond = 'abs(am) > 22 & abs(am) < 28',
  action = 'stop',
  plot = TRUE))
# NB: clicks are often introduced at the onset and offset - fade and renormalize
s_filt = s_filt / max(abs(s_filt))
# playme(s_filt, samplingRate = 16000)
spectrogram(s_filt, 16000, windowLength = 25)

# Using filterMS()
ms = modulationSpectrum(s, 16000, returnComplex = TRUE)$detailed$complex

# Remove all AM over 25 Hz
filterMS(ms, amCond = 'abs(am) > 25', protect = FALSE)
filterMS(ms, amCond = 'abs(am) > 25', protect = TRUE)
filterMS(ms, amCond = 'abs(am) > 25', protect = TRUE, smooth = NA)

# amCond and fmCond are OR-conditions
filterMS(ms, amCond = 'abs(am) > 15', fmCond = 'abs(fm) > 5', action = 'stop')
filterMS(ms, amCond = 'abs(am) > 15', fmCond = 'abs(fm) > 5', action = 'pass')
filterMS(ms, amCond = 'abs(am) > 10 & abs(am) < 25', action = 'stop')

# jointCond is more flexible
filterMS(ms, jointCond = 'am * fm < 5', action = 'stop')
filterMS(ms, jointCond = 'am^2 + (fm*3)^2 < 200', action = 'pass')

# So:
filterMS(ms, jointCond = 'abs(am) > 5 | abs(fm) < 5')  # general
# ...is the same as:
filterMS(ms, amCond = 'abs(am) > 5', fmCond = 'abs(fm) < 5')  # faster

## Not run: 
data('speechEx', package = 'soundgen')
samplingRate = speechEx@samp.rate
playme(speechEx)
spectrogram(speechEx)

# Remove AM above 3 Hz from a bit of speech (removes most temporal details)
s_filt1 = fade(filterSoundByMS(speechEx, amCond = 'abs(am) > 3',
                          action = 'stop',
                          smooth = c(0.5, 2)))
spectrogram(s_filt1, samplingRate = samplingRate)
playme(s_filt1, samplingRate)

# Remove FM above 3 cycles/kHz (masks f0 but preserves formants)
s_filt3 = fade(filterSoundByMS(speechEx,
  fmCond = 'abs(fm) > 3', action = 'stop',
  logSpec = TRUE, smooth = c(1, 5),
  protect = c('level', 'envelope')))
spectrogram(s_filt3, samplingRate = samplingRate)
playme(s_filt3, samplingRate)

# Joint spectro-temporal filtering with a symmetric condition
s_filt4 = fade(filterSoundByMS(speechEx,
  jointCond = 'am^2 + (fm * 3)^2 < 300', action = 'stop',
  logSpec = TRUE, dynamicRange = 60, smooth = c(1, 10),
  protect = c('level', 'meanSpectrum')))
spectrogram(s_filt4, samplingRate = samplingRate)
playme(s_filt4, samplingRate)

# Process all files in a folder, save filtered audio and plots
s_filt = filterSoundByMS('~/Downloads/temp2',
  saveAudio = TRUE, savePlots = TRUE,
  amCond = 'abs(am) > 15', fmCond = 'abs(fm) > 5',
  action = 'stop', smooth = c(0.5, 5),
  nIter = 10)

## End(Not run)

soundgen documentation built on Sept. 20, 2026, 5:07 p.m.