View source: R/compareSounds.R
| compareSounds | R Documentation |
Computes distances between sounds based on comparing their spectrogram-like
representations. compareSounds takes two sounds or feature matrices as
input, whereas compareFolder takes a path to a folder with audio files
or a list of feature matrices and returns a matrix of pairwise distances
between them. Feature matrices are normalized and compared with
Dynamic Time Warp (DTW), correlation, cosine distance, or pixel by pixel.
compareSounds(
x,
y,
samplingRate = NULL,
specFun = "melspec",
specFun_pars = list(),
logSpec = FALSE,
method = c("cor", "cosine", "diff", "dtw"),
padWith = NA,
padDir = c("central", "left", "right"),
dtw_pars = list()
)
compareFolder(
myfolder = NULL,
spectrograms = NULL,
matchAllLengths = FALSE,
specFun = "melspec",
specFun_pars = list(),
logSpec = FALSE,
method = c("cor", "cosine", "diff", "dtw"),
padWith = NA,
padDir = c("central", "left", "right"),
dtw_pars = list(),
cores = 1,
reportEvery = NULL
)
x, y |
either two matrices (spectrograms or feature matrices) or two sounds to be compared (numeric vectors, Wave objects, or paths to wav/mp3 files) |
samplingRate |
if one or both inputs are numeric vectors, specify sampling rate, Hz. This does not resample audio. For meaningful comparisons, audio inputs should already have the same sampling rate |
specFun |
the function used to extract a spectrogram-like feature matrix. Can be a string or a custom function that takes audio (numeric vector) as the first argument and returns a spectrogram-like matrix with time in columns and features in rows (see examples). Supported strings:
|
specFun_pars |
a list of parameters passed to |
logSpec |
if TRUE, applies a log transform to the spectrograms before normalization |
method |
method(s) of comparing spectrograms of two sounds:
"cor" = Pearson's correlation distance; "cosine" = cosine distance;
"diff" = normalized absolute difference; "dtw" = multivariate Dynamic Time
Warp with |
padWith |
if the durations of x and y are not identical, the compared
spectrograms are either padded with silence ( |
padDir |
if padding, specify where to add zeros or NAs: before the sound ('left'), after the sound ('right'), or on both sides ('central') |
dtw_pars |
a list of parameters passed to |
myfolder |
path to folder containing audio files to compare |
spectrograms |
a list of spectrogram-like feature matrices to use
instead of analyzing the audio - extracted acoustic features,
modulation spectra, similarity matrices, ... (overrides |
matchAllLengths |
if TRUE, all spectrograms are length-matched - e.g.,
if 100 sounds are compared and |
cores |
number of cores for parallel processing |
reportEvery |
when processing multiple inputs, report estimated time
left every |
If the input is audio, several methods of producing spectrograms are
available ("specFun"). For more customized options, just prepare your
spectrograms or feature matrices first (time in columns, features like pitch,
peak frequency, etc. in rows), and then pass them to compareSounds
(see examples). All methods except for DTW require that the compared matrices
should be of the same size. Compared sounds should ideally have the same
sampling rate. If they differ, row (frequency bin) truncation is performed by
position, keeping only the first min(nrow) rows, which approximates keeping
frequencies up to the lower Nyquist frequency when both spectrograms use the
same frequency resolution. In case of differences in duration, the shorter
sound is padded with 0 (silence) or NA, as controlled by arguments
padWith, padDir. If passing custom feature matrices, ensure that they
have the same dimensions or that padding with 0 (silence) makes sense.
compareSounds returns a dataframe with two columns: "method"
for the method(s) used, and "distance" for the distance between the two
sounds calculated with that method. The range of distances is [0, 1]
for "cor", "cosine", and "diff", and [0, Inf) for "dtw".
compareFolder returns a list of distance matrices (dist
objects), one for each method.
s1 = soundgen(sylLen = 100, pitch = c(80, 180), formants = 'a')
s2 = soundgen(sylLen = 120, pitch = c(150, 350), formants = 'u')
compareSounds(s1, s2, samplingRate = 16000, method = c('cor', 'cosine', 'diff'))
# spectrogram(s1); playme(s1)
# spectrogram(s2); playme(s2)
## Not run:
# NB: install the "dtw" library to run the examples
# compare all sounds in a folder, e.g.:
target = '~/Documents/Research/zz_test_audio/temp_long'
cf = compareFolder(target, logSpec = TRUE)
mds = as.data.frame(cmdscale(cf$cor))
plot(mds, type = 'n'); text(mds, labels = abbreviate(rownames(mds)))
# or use manually produced spectrograms
sp = spectrogram(target, windowLength = c(10, 40), overlap = 75,
yScale = 'ERB', output = 'processed', plot = FALSE, cores = 4)
image(sp[[1]])
cf1 = compareFolder(spectrograms = sp)
mds1 = as.data.frame(cmdscale(cf1$cor))
plot(mds1, type = 'n'); text(mds1, labels = abbreviate(rownames(mds1)))
# extract a spectrogram-like representation using a custom function
# (e.g., full-resolution analytic envelopes instead of downsampled RMS)
compareSounds(s1, s2, samplingRate = 16000,
specFun = function(x) matrix(hilbert_approx(x)$envelope, nrow = 1))
# some more examples
s1 = soundgen(formants = 'a', play = TRUE)
s2 = soundgen(formants = 'ae', play = TRUE)
s3 = soundgen(formants = 'eae', sylLen = 700, play = TRUE)
s4 = runif(8000, -1, 1) # white noise
compareSounds(s1, s2, samplingRate = 16000)
compareSounds(s1, s4, samplingRate = 16000)
# the central section of s3 is more similar to s1 than is the beg/end of s3
compareSounds(s1, s3, samplingRate = 16000, padDir = 'left')
compareSounds(s1, s3, samplingRate = 16000, padDir = 'central')
# padding with 0 penalizes differences in duration, whereas padding with NA
# is like saying we only care about the overlapping part
compareSounds(s1, s4, samplingRate = 16000, padWith = 0)
compareSounds(s1, s4, samplingRate = 16000, padWith = NA)
# different types of spectrograms produce quite different results
compareSounds(s1, s3, samplingRate = 16000, specFun = 'stft')
compareSounds(s1, s3, samplingRate = 16000, specFun = 'melspec')
compareSounds(s1, s3, samplingRate = 16000, specFun = 'mfcc')
compareSounds(s1, s3, samplingRate = 16000, specFun = 'audSpec')
# pass additional control parameters to specFun and DTW
compareSounds(s1, s3, samplingRate = 16000,
specFun = 'melspec',
specFun_pars = list(nbands = 128),
dtw_pars = list(dist.method = "Manhattan"))
# use feature matrices instead of spectrograms
# (time in columns, features in rows)
a1 = t(as.matrix(analyze(s1, samplingRate = 16000)$detailed))
a1 = a1[4:nrow(a1), ]; a1[is.na(a1)] = 0 # don't use dur and time stamps
a2 = t(as.matrix(analyze(s2, samplingRate = 16000)$detailed))
a2 = a2[4:nrow(a2), ]; a2[is.na(a2)] = 0
a4 = t(as.matrix(analyze(s4, samplingRate = 16000)$detailed))
a4 = a4[4:nrow(a4), ]; a4[is.na(a4)] = 0
compareSounds(a1, a2, method = c('cosine', 'dtw'))
compareSounds(a1, a4, method = c('cosine', 'dtw'))
## End(Not run)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.