| estimate_jsd | R Documentation |
Use this function when Jensen-Shannon divergence (JSD) or Jensen-Shannon
distance is the primary outcome. If you want to compare JSD with Pillai,
Bhattacharyya, Mahalanobis, and percent-overlap metrics, start with
phontrast() instead.
estimate_jsd(
data,
features,
category_col,
group_col = NULL,
do_boot = FALSE,
n_boot = 1000,
min_tokens = 20,
est_distance = FALSE,
conf_level = 0.95,
bw = c("Hpi", "Hscv", "Hpi.diag", "scott.diag"),
eval_on = c("pooled", "group1", "group2", "pooled_sample"),
eval_n = NULL,
eval_seed = NULL,
engine = c("ks", "fast_diag", "fast_diagonal"),
chunk_size = 1000L,
method = c("mc", "legacy"),
density = c("kde", "mvnorm"),
mc_n = 10000L,
bw_scale = 1,
...
)
data |
Data frame with at least |
features |
Character vector of feature column names (e.g., c("F1","F2")). |
category_col |
Name of the column giving the two-way category factor. |
group_col |
Optional character vector of one or more grouping columns.
If provided, returns per-group JSD. Multiple grouping columns are combined
into a labeled |
do_boot |
Logical; if TRUE, run nonparametric bootstrap. |
n_boot |
Number of bootstrap resamples. |
min_tokens |
Minimum total tokens required (globally or per group). |
est_distance |
Logical; if TRUE, return Jensen-Shannon distance (sqrt of divergence). |
conf_level |
Confidence level for bootstrap interval. |
bw |
Bandwidth selection method passed to |
eval_on |
KDE evaluation points passed to |
eval_n |
Optional maximum number of KDE evaluation points. |
eval_seed |
Optional integer seed for KDE evaluation-point subsampling. |
engine |
KDE evaluation engine passed to |
chunk_size |
Chunk size for |
method |
Estimator passed to |
density |
Density model behind the estimate, passed to
|
mc_n |
Positive integer; number of Monte-Carlo samples drawn from each
fitted Gaussian when |
bw_scale |
Positive number multiplying the selected kernel bandwidth on
the standard-deviation scale (default |
... |
Additional arguments passed to |
JSD is bounded from 0 to 1 for two equally weighted distributions. Larger
values indicate greater category separation. Jensen-Shannon distance is
sqrt(JSD) and has the same direction of interpretation.
A tibble. Global: one row with columns
scope, n_tokens, n_boot, conf_level, jsd_point, jsd_mean, jsd_sd,
ci_lower, ci_upper, jsd_low, jsd_high.
Grouped: one row per group with columns
scope, group, n_tokens, n_boot, conf_level, jsd_point, jsd_mean,
jsd_sd, ci_lower, ci_upper, jsd_low, jsd_high.
ci_lower and ci_upper are the preferred confidence interval
columns; jsd_low and jsd_high are retained as legacy aliases.
set.seed(2026)
vowels <- data.frame(
speaker = rep(c("s01", "s02"), each = 60),
vowel = rep(rep(c("ih", "eh"), each = 30), 2),
f1 = c(
rnorm(30, 500, 55), rnorm(30, 560, 60),
rnorm(30, 510, 60), rnorm(30, 575, 65)
),
f2 = c(
rnorm(30, 1980, 150), rnorm(30, 1880, 155),
rnorm(30, 1960, 160), rnorm(30, 1840, 165)
)
)
# Point estimate of JSD (fast), globally and by speaker.
estimate_jsd(
data = vowels,
features = c("f1", "f2"),
category_col = "vowel"
)
estimate_jsd(vowels, c("f1", "f2"), "vowel", group_col = "speaker")
# Bootstrap confidence intervals, shown on a single feature: multivariate
# bootstraps work the same way but repeat multivariate bandwidth selection
# on every resample, so they take correspondingly longer. Increase n_boot
# for real analyses.
estimate_jsd(vowels, "f1", "vowel", do_boot = TRUE, n_boot = 20)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.