| ssm | R Documentation |
Calculates the self-similarity matrix and novelty vector of a sound. The self-similarity matrix is produced by comparing all pairs of frames in the input sound. Novelty is calculated by convolving the self-similarity matrix with a tapered checkerboard kernel. The positive lobes of the kernel represent coherence (self-similarity within the regions on either side of the center point) and the negative lobes anti-coherence (cross-similarity between these two regions). Since novelty is the dot product of the checkerboard kernel with the SSM, it is high when the two regions are self-similar (internally consistent) but different from each other.
ssm(
x,
samplingRate = NULL,
from = NULL,
to = NULL,
specFun = "melspec",
specFun_pars = list(),
logSpec = FALSE,
normalize = FALSE,
simil = c("cosine", "cor"),
kernelLen = 1000,
kernelSD = 0.5,
padWith = 0,
ssmWin = 1,
summaryFun = c("mean", "sd"),
output = c("ssm", "novelty"),
reportEvery = NULL,
cores = 1,
plot = TRUE,
savePlots = FALSE,
embed = FALSE,
main = NULL,
heights = c(2, 1),
width = 900,
height = 500,
units = "px",
res = NA,
specPars = list(colorTheme = c("bw", "seewave", "heat.colors", "...")[2], xlab =
"Time"),
ssmPars = list(colorTheme = c("bw", "seewave", "heat.colors", "...")[2], xlab = "Time",
ylab = "Time"),
noveltyPars = list(type = "l", col = "black", lwd = 1)
)
x |
path to a folder, one or more wav or mp3 files c('file1.wav', 'file2.mp3'), Wave object, numeric vector, or a list of Wave objects or numeric vectors |
samplingRate |
sampling rate of |
from, to |
if NULL (default), analyzes the whole sound, otherwise from...to (s) |
specFun |
the function used to extract a spectrogram-like feature matrix. Can be a string or a custom function that takes audio (numeric vector) as the first argument and returns a spectrogram-like matrix with time in columns and features in rows (see examples). A precomputed matrix is also accepted (features in rows, time [ms] in columns, numeric rownames for plotting). Supported strings:
|
specFun_pars |
a list of parameters passed to |
logSpec |
if TRUE, the input is log-transformed prior to calculating self-similarity |
normalize |
if TRUE, the spectrum of each STFT frame (or each column of feature matrix) is normalized to the same range prior to calculating self-similarity |
simil |
method for comparing frames: "cosine" = cosine similarity, "cor" = Pearson's correlation |
kernelLen |
length of checkerboard kernel for calculating novelty, ms (larger values favor global, slow vs. local, fast novelty) |
kernelSD |
SD of checkerboard kernel evaluated over [-1, 1]: for ex., if kernelSD = 0.5, the kernel spans approximately ±2 SDs |
padWith |
how to treat edges when calculating novelty: NA = pad with NA (ignores edges in correlation), 0 = pad with zeros |
ssmWin |
window for averaging SSM, frames (has a smoothing effect and speeds up the processing) |
summaryFun |
functions used to summarize novelty, eg |
output |
what to include in |
reportEvery |
when processing multiple inputs, report estimated time
left every |
cores |
number of cores for parallel processing |
plot |
if TRUE, plots the SSM |
savePlots |
if TRUE, creates a subdirectory in the input directory (if input is a file or folder) or in the working directory (if input is a vector etc), named after the function (eg "spectrogram/"). All plots and audio files (if any) are saved in this new directory. If there are multiple inputs, an html notebook is also created for easy viewing and listening |
embed |
if TRUE and savePlots is set and there are multiple inputs, all saved images and audio (if any) are embedded in the exported html notebook for easy sharing; if FALSE, the html file links to separate images and audio files (but separate files are still saved). NB: for this to work, package "base64enc" must be installed |
main |
plot title |
heights |
relative sizes of the SSM and spectrogram/novelty plot |
width, height, units, res |
graphical parameters for saving plots passed to
|
specPars |
graphical parameters passed to |
ssmPars |
graphical parameters passed to |
noveltyPars |
graphical parameters passed to
|
A list with two top-level elements: $detailed and
$summary.
$detailed contains per-file results selected with the output
argument. If multiple sounds are analyzed, $detailed is a named list
of per-sound lists. If a single sound is analyzed, it is simplified to a
single list. Each list may contain:
self-similarity matrix
novelty vector
$summary is a dataframe with summaries of novelty per sound (one
row per file), or NULL if summaryFun = NULL.
Foote, J. (1999, October). Visualizing music and audio using self-similarity. In Proceedings of the seventh ACM international conference on Multimedia (Part 1) (pp. 77-80). ACM.
Foote, J. (2000). Automatic audio segmentation using a measure of audio novelty. In Multimedia and Expo, 2000. ICME 2000. 2000 IEEE International Conference on (Vol. 1, pp. 452-455). IEEE.
spectrogram modulationSpectrum
segment
sound = c(soundgen(),
soundgen(nSyl = 4, sylLen = 50, pauseLen = 70,
formants = NA, pitch = c(500, 330)))
# playme(sound)
# detailed, local features (captures each syllable)
s1 = ssm(sound, samplingRate = 16000, kernelLen = 100)
# more global features (captures the transition b/w the two sounds)
s2 = ssm(sound, samplingRate = 16000, kernelLen = 400)
s2$summary
s2$detailed$novelty # novelty contour
## Not run:
ssm(sound, samplingRate = 16000,
specFun = 'mfcc', simil = 'cor', normalize = TRUE,
ssmWin = 10, # speed up the processing
kernelLen = 300, # global features
specPars = list(colorTheme = 'seewave'),
ssmPars = list(col = rainbow(100)),
noveltyPars = list(type = 'l', lty = 3, lwd = 2))
# Custom input: produce a nice spectrogram first, then feed it into ssm()
sp = spectrogram(sound, 16000, windowLength = c(5, 40), contrast = .3,
output = 'processed') # return the modified spectrogram
ssm(sound, 16000, kernelLen = 400, specFun = sp)
# Custom input: use acoustic features returned by analyze()
an = analyze(sound, 16000, windowLength = 20, novelty = NULL)
feature_mat = t(an$detailed[, 4:ncol(an$detailed)]) # or select pitch, HNR, ...
feature_mat = t(apply(feature_mat, 1, scale)) # z-transform all variables
feature_mat[is.na(feature_mat)] = 0 # get rid of NAs
colnames(feature_mat) = an$detailed$time # time stamps in ms
rownames(feature_mat) = 1:nrow(feature_mat)
image(t(feature_mat)) # not a spectrogram, just a feature matrix
ssm(sound, 16000, kernelLen = 500, specFun = feature_mat, logSpec = FALSE,
specPars = list(ylab = 'Feature'))
## End(Not run)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.