getLoudness: Get loudness

View source: R/loudness.R

getLoudnessR Documentation

Get loudness

Description

Estimates subjective loudness and sharpness of audio. Based on EMBSD speech quality measure, particularly the MATLAB code in Yang (1999) and Timoney et al. (2004). Note that there are many ways to estimate loudness and many other factors, ignored by this model, that could influence subjectively experienced loudness. Please treat the output with a healthy dose of skepticism! Also note that the absolute value of calculated loudness critically depends on the chosen "measured" sound pressure level (SPL). getLoudness estimates how loud a sound will be experienced if it is played back at an SPL of SPL_measured dB. The most meaningful way to use the output is to compare the loudness of several sounds analyzed with identical settings or of different segments within the same recording.

Usage

getLoudness(
  x,
  samplingRate = NULL,
  scale = NULL,
  from = NULL,
  to = NULL,
  input = c("spec", "audSpec"),
  windowLength = 50,
  step = NULL,
  overlap = 50,
  SPL_measured = 70,
  spreadSpectrum = FALSE,
  sharpnessMethod = c("aures", "DIN45692", "bismarck"),
  summaryFun = c("mean", "median", "sd"),
  reportEvery = NULL,
  cores = 1,
  plot = TRUE,
  savePlots = FALSE,
  embed = FALSE,
  main = NULL,
  ylim = NULL,
  width = 900,
  height = 500,
  units = "px",
  res = NA,
  mar = c(5.1, 4.1, 4.1, 4.1),
  ...
)

Arguments

x

path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors

samplingRate

sampling rate of x (only needed if x is a numeric vector)

scale

maximum possible amplitude of input, used to normalize the input vector (only needed if x is a numeric vector)

from, to

if NULL (default), analyzes the whole sound, otherwise from...to (s)

input

"spec" = power spectrogram warped to bark scale with audspec, "audSpec" = auditory spectrogram produced by convolving the signal with a bank of gammatone filters using audSpectrogram (much slower, but more physiologically accurate)

windowLength

length of the analysis window, ms

step

step between successive windows, ms; if provided, overrides overlap; because digital audio is sampled at discrete time intervals of 1/samplingRate, the actual step and thus the time stamps of STFT frames may be slightly different - e.g., 24.98866 instead of 25.0 ms

overlap

overlap between successive windows, %

SPL_measured

sound pressure level at which the sound is presented relative to some reference (the conventional threshold is 2e-5 Pa), dB

spreadSpectrum

if TRUE, applies a spreading function to account for frequency masking

sharpnessMethod

the method of calculating sharpness (mostly differ in weighting functions; only "aures" depends on SPL)

summaryFun

functions used to summarize each acoustic characteristic, eg c('mean', 'sd'); user-defined functions are fine (see examples); NAs are omitted automatically for mean/median/sd/min/max/range/sum, otherwise take care of NAs yourself

reportEvery

when processing multiple inputs, report estimated time left every reportEvery iterations (NULL = default, NA = don't report); see reportTime

cores

number of cores for parallel processing

plot

if TRUE, produces a plot of the results

savePlots

if TRUE, creates a subdirectory in the input directory (if input is a file or folder) or in the working directory (if input is a vector etc), named after the function (eg "spectrogram/"). All plots and audio files (if any) are saved in this new directory. If there are multiple inputs, an html notebook is also created for easy viewing and listening

embed

if TRUE and savePlots is set and there are multiple inputs, all saved images and audio (if any) are embedded in the exported html notebook for easy sharing; if FALSE, the html file links to separate images and audio files (but separate files are still saved). NB: for this to work, package "base64enc" must be installed

main

plot title

ylim

frequency range to plot, kHz (defaults to 0 to Nyquist frequency). NB: still in kHz, even if yScale = bark, mel, or ERB

width, height, units, res

graphical parameters for saving plots passed to png

mar

margins of the spectrogram

...

other plotting parameters passed to spectrogram

Details

Algorithm: calibrates the sound to the desired SPL (Timoney et al., 2004), extracts a spectrogram with frequencies on the bark scale, optionally spreads the spectrum to account for frequency masking across the critical bands (Yang, 1999), converts dB to phon by using standard equal loudness curves (ISO 226), converts phon to sone (Timoney et al., 2004), sums across all critical bands, and applies a correction coefficient to standardize output. Calibrated so as to return a loudness of 1 sone for a 1 kHz pure tone with SPL of 40 dB and spreadSpectrum = FALSE. Sharpness is calculated as the weighted first moment of specific loudness on the Bark scale.

Value

A list with two top-level elements: $detailed and $summary.

$detailed contains per-file results. If multiple sounds are analyzed, $detailed is a list of per-sound lists. If a single sound is analyzed, it is simplified to a single list. Each list contains:

loudness

a vector of loudness in sone units per STFT frame

specSone

spectrum in bark-sone: a matrix of loudness values in sone, with frequency on the bark scale in rows and time (STFT frames) in columns

loudnessPhon, specPhon

same in phon instead of sone units

sharpness

a vector of sharpness in acum units per STFT frame

audSpec

auditory spectrogram used to calculate loudness and sharpness

$summary is a dataframe of summary loudness measures (one row per file). If summaryFun is NULL, $summary is NULL.

References

  • ISO 226 as implemented by Jeff Tackett (2005) on https://www.mathworks.com/matlabcentral/fileexchange/ 7028-iso-226-equal-loudness-level-contour-signal

  • Timoney, J., Lysaght, T., Schoenwiesner, M., & MacManus, L. (2004). Implementing loudness models in matlab.

  • Yang, W. (1999). Enhanced Modified Bark Spectral Distortion (EMBSD): An Objective Speech Quality Measure Based on Audible Distortion and Cognitive Model. Temple University.

See Also

getRMS analyze

Examples

sounds = list(
  noise_1KHz = soundgen:::zeroOne(bandpass(rnorm(8000), 16000,
    lwr = 900, upr = 1100)) * 2 -1, # narrow-band noise at 1 KHz
  white_noise = runif(8000, -1, 1),  # white noise
  white_noise2 = runif(8000, -1/2, 1/2),  # ~6 dB quieter
  pure_tone_1KHz = sin(2*pi*1000/16000*(1:8000)),  # pure tone at 1 kHz
  pure_tone_100Hz = sin(2*pi*100/16000*(1:8000))  # pure tone at 100 Hz
)
# playme(sounds)
l = getLoudness(
    x = sounds, samplingRate = 16000, scale = 1,
    windowLength = 40, step = NULL, input = c('spec', 'audSpec')[1],
    overlap = 50, SPL_measured = 60,
    plot = FALSE)
l$summary
# loudness depends on amplitude if "scale" is provided (cf. sounds 2 and 3)
# narrowband noise / tone at 1 kHz, 60 dB: sharpness ~=1 acum, loudness ~=4 sone

# a steady glissando from 125 to 8000 Hz (constant on a musical scale)
pitch = exp(seq(log(125), log(8000), length.out = 16000))
s = sinpi(2 * cumsum(pitch) / 16000)
l1 = getLoudness(s, samplingRate = 16000, SPL_measured = 70)
# steady SPL, but variable loudness

# The estimated loudness and sharpness depend on target SPL
l2 = getLoudness(s, samplingRate = 16000, SPL_measured = 40, plot = FALSE)
l1$summary$loudness_mean
l2$summary$loudness_mean

# ...but not (much) on windowLength and samplingRate
l3 = getLoudness(s, samplingRate = 16000, SPL_measured = 40,
  windowLength = 50, plot = FALSE)
l3$summary$loudness_mean

## Not run: 
# Using auditory spectrogram as input instead of STFT (slower)
l4 = getLoudness(s, samplingRate = 16000, SPL_measured = 40, input = 'audSpec')
l4$summary$loudness_mean

# Process all audio files in a folder
l5 = getLoudness('~/Downloads/temp', savePlots = TRUE)
l5$summary

## End(Not run)

soundgen documentation built on Sept. 20, 2026, 5:07 p.m.