| global_boot_jsd | R Documentation |
Computes a single Jensen-Shannon divergence (JSD) value for two categories in an n-dimensional acoustic space, together with bootstrap-based confidence intervals obtained by resampling tokens with replacement.
global_boot_jsd(
data,
features,
category_col,
n_boot = 1000,
min_tokens = 20,
est_distance = FALSE,
conf_level = 0.95,
bw = c("Hpi", "Hscv", "Hpi.diag", "scott.diag"),
eval_on = c("pooled", "group1", "group2", "pooled_sample"),
eval_n = NULL,
eval_seed = NULL,
engine = c("ks", "fast_diag", "fast_diagonal"),
chunk_size = 1000L,
method = c("mc", "legacy"),
density = c("kde", "mvnorm"),
mc_n = 10000L,
...
)
data |
Data frame containing at least the category column and the feature columns. |
features |
Character vector of column names giving the acoustic dimensions (e.g., c("f1", "f2") or paste0("mfcc", 1:13)). |
category_col |
String; name of the column giving the two categories to compare (e.g., "vowel"). Must have exactly two unique values. |
n_boot |
Integer; number of bootstrap resamples. |
min_tokens |
Minimum total number of non-missing tokens required. |
est_distance |
Logical; if TRUE, return Jensen-Shannon distance (sqrt of divergence) instead of divergence. |
conf_level |
Confidence level for bootstrap intervals. |
bw |
Bandwidth selection method passed to |
eval_on |
KDE evaluation points passed to |
eval_n |
Optional maximum number of KDE evaluation points. |
eval_seed |
Optional integer seed for KDE evaluation-point subsampling. |
engine |
KDE evaluation engine passed to |
chunk_size |
Chunk size for |
method |
Estimator passed to |
density |
Density model passed to |
mc_n |
Positive integer; number of Monte-Carlo samples drawn from each
fitted Gaussian when |
... |
Additional arguments passed to |
This is the "group-wise" version of JSD: it ignores speakers and treats all tokens as coming from a single population for each category.
A one-row data frame with columns:
n_tokens - total number of tokens used
n_boot - number of successful bootstrap samples
conf_level - confidence level used for the interval
jsd_point - JSD on the full dataset
jsd_mean - mean JSD across bootstrap samples
jsd_sd - standard deviation of bootstrap JSD
ci_lower, ci_upper - bootstrap confidence interval
jsd_low, jsd_high - legacy aliases for
ci_lower and ci_upper
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.