global_boot_jsd: Global JSD with bootstrap confidence interval

View source: R/global_jsd.R

global_boot_jsdR Documentation

Global JSD with bootstrap confidence interval

Description

Computes a single Jensen-Shannon divergence (JSD) value for two categories in an n-dimensional acoustic space, together with bootstrap-based confidence intervals obtained by resampling tokens with replacement.

Usage

global_boot_jsd(
  data,
  features,
  category_col,
  n_boot = 1000,
  min_tokens = 20,
  est_distance = FALSE,
  conf_level = 0.95,
  bw = c("Hpi", "Hscv", "Hpi.diag", "scott.diag"),
  eval_on = c("pooled", "group1", "group2", "pooled_sample"),
  eval_n = NULL,
  eval_seed = NULL,
  engine = c("ks", "fast_diag", "fast_diagonal"),
  chunk_size = 1000L,
  method = c("mc", "legacy"),
  density = c("kde", "mvnorm"),
  mc_n = 10000L,
  ...
)

Arguments

data

Data frame containing at least the category column and the feature columns.

features

Character vector of column names giving the acoustic dimensions (e.g., c("f1", "f2") or paste0("mfcc", 1:13)).

category_col

String; name of the column giving the two categories to compare (e.g., "vowel"). Must have exactly two unique values.

n_boot

Integer; number of bootstrap resamples.

min_tokens

Minimum total number of non-missing tokens required.

est_distance

Logical; if TRUE, return Jensen-Shannon distance (sqrt of divergence) instead of divergence.

conf_level

Confidence level for bootstrap intervals.

bw

Bandwidth selection method passed to jsd_kde_nd().

eval_on

KDE evaluation points passed to jsd_kde_nd().

eval_n

Optional maximum number of KDE evaluation points.

eval_seed

Optional integer seed for KDE evaluation-point subsampling.

engine

KDE evaluation engine passed to jsd_kde_nd(). "fast_diagonal" is accepted as an alias for "fast_diag".

chunk_size

Chunk size for engine = "fast_diag".

method

Estimator passed to jsd_kde_nd(): "mc" (default) or "legacy" (pre-1.2.0 self-normalized estimate). Ignored when density = "mvnorm".

density

Density model passed to estimate_jsd(): "kde" (default) or "mvnorm" (fit one multivariate normal per category and estimate JSD between the two Gaussians by Monte-Carlo).

mc_n

Positive integer; number of Monte-Carlo samples drawn from each fitted Gaussian when density = "mvnorm" (default 10000). Ignored when density = "kde".

...

Additional arguments passed to jsd_kde_nd().

Details

This is the "group-wise" version of JSD: it ignores speakers and treats all tokens as coming from a single population for each category.

Value

A one-row data frame with columns:

  • n_tokens - total number of tokens used

  • n_boot - number of successful bootstrap samples

  • conf_level - confidence level used for the interval

  • jsd_point - JSD on the full dataset

  • jsd_mean - mean JSD across bootstrap samples

  • jsd_sd - standard deviation of bootstrap JSD

  • ci_lower, ci_upper - bootstrap confidence interval

  • jsd_low, jsd_high - legacy aliases for ci_lower and ci_upper


phontrast documentation built on Oct. 7, 2026, 5:06 p.m.