estimate_jsd: Estimate Jensen-Shannon divergence or distance between two...

View source: R/jsd_wrappers.R

estimate_jsdR Documentation

Estimate Jensen-Shannon divergence or distance between two categories

Description

Use this function when Jensen-Shannon divergence (JSD) or Jensen-Shannon distance is the primary outcome. If you want to compare JSD with Pillai, Bhattacharyya, Mahalanobis, and percent-overlap metrics, start with phontrast() instead.

Usage

estimate_jsd(
  data,
  features,
  category_col,
  group_col = NULL,
  do_boot = FALSE,
  n_boot = 1000,
  min_tokens = 20,
  est_distance = FALSE,
  conf_level = 0.95,
  bw = c("Hpi", "Hscv", "Hpi.diag", "scott.diag"),
  eval_on = c("pooled", "group1", "group2", "pooled_sample"),
  eval_n = NULL,
  eval_seed = NULL,
  engine = c("ks", "fast_diag", "fast_diagonal"),
  chunk_size = 1000L,
  method = c("mc", "legacy"),
  density = c("kde", "mvnorm"),
  mc_n = 10000L,
  bw_scale = 1,
  ...
)

Arguments

data

Data frame with at least category_col and features.

features

Character vector of feature column names (e.g., c("F1","F2")).

category_col

Name of the column giving the two-way category factor.

group_col

Optional character vector of one or more grouping columns. If provided, returns per-group JSD. Multiple grouping columns are combined into a labeled group value such as "Sex=F | Style=read".

do_boot

Logical; if TRUE, run nonparametric bootstrap.

n_boot

Number of bootstrap resamples.

min_tokens

Minimum total tokens required (globally or per group).

est_distance

Logical; if TRUE, return Jensen-Shannon distance (sqrt of divergence).

conf_level

Confidence level for bootstrap interval.

bw

Bandwidth selection method passed to jsd_kde_nd().

eval_on

KDE evaluation points passed to jsd_kde_nd().

eval_n

Optional maximum number of KDE evaluation points.

eval_seed

Optional integer seed for KDE evaluation-point subsampling.

engine

KDE evaluation engine passed to jsd_kde_nd(). "fast_diagonal" is accepted as an alias for "fast_diag".

chunk_size

Chunk size for engine = "fast_diag".

method

Estimator passed to jsd_kde_nd(): "mc" (default) for the Monte-Carlo plug-in estimate of the continuous JSD, or "legacy" to reproduce the pre-1.2.0 self-normalized sample-point estimate. Ignored when density = "mvnorm".

density

Density model behind the estimate, passed to jsd_kde_nd(): "kde" (default) estimates each category's density by kernel density estimation; "mvnorm" fits one multivariate normal per category and estimates the continuous JSD between the two Gaussians by Monte-Carlo. Under "mvnorm" the KDE-specific arguments do not apply and the Monte-Carlo sample size is set by mc_n.

mc_n

Positive integer; number of Monte-Carlo samples drawn from each fitted Gaussian when density = "mvnorm" (default 10000). Ignored when density = "kde".

bw_scale

Positive number multiplying the selected kernel bandwidth on the standard-deviation scale (default 1); 0.5 and 2 give the halved and doubled bandwidths of the smoothing-sensitivity check. Passed to jsd_kde_nd(); ignored when density = "mvnorm".

...

Additional arguments passed to jsd_kde_nd() (e.g., loo).

Details

JSD is bounded from 0 to 1 for two equally weighted distributions. Larger values indicate greater category separation. Jensen-Shannon distance is sqrt(JSD) and has the same direction of interpretation.

Value

A tibble. Global: one row with columns scope, n_tokens, n_boot, conf_level, jsd_point, jsd_mean, jsd_sd, ci_lower, ci_upper, jsd_low, jsd_high. Grouped: one row per group with columns scope, group, n_tokens, n_boot, conf_level, jsd_point, jsd_mean, jsd_sd, ci_lower, ci_upper, jsd_low, jsd_high. ci_lower and ci_upper are the preferred confidence interval columns; jsd_low and jsd_high are retained as legacy aliases.

Examples

set.seed(2026)
vowels <- data.frame(
  speaker = rep(c("s01", "s02"), each = 60),
  vowel = rep(rep(c("ih", "eh"), each = 30), 2),
  f1 = c(
    rnorm(30, 500, 55), rnorm(30, 560, 60),
    rnorm(30, 510, 60), rnorm(30, 575, 65)
  ),
  f2 = c(
    rnorm(30, 1980, 150), rnorm(30, 1880, 155),
    rnorm(30, 1960, 160), rnorm(30, 1840, 165)
  )
)

# Point estimate of JSD (fast), globally and by speaker.
estimate_jsd(
  data = vowels,
  features = c("f1", "f2"),
  category_col = "vowel"
)
estimate_jsd(vowels, c("f1", "f2"), "vowel", group_col = "speaker")

# Bootstrap confidence intervals, shown on a single feature: multivariate
# bootstraps work the same way but repeat multivariate bandwidth selection
# on every resample, so they take correspondingly longer. Increase n_boot
# for real analyses.
estimate_jsd(vowels, "f1", "vowel", do_boot = TRUE, n_boot = 20)

phontrast documentation built on Oct. 7, 2026, 5:06 p.m.