phontrast: Compute and compare phonological contrast metrics

View source: R/compare_overlap_metrics.R

phontrastR Documentation

Compute and compare phonological contrast metrics

Description

phontrast() is the package's main entry point. It computes one or more category separation and overlap metrics for a two-category phonological contrast in a single call: Jensen-Shannon divergence and distance, the Pillai-Bartlett trace, Bhattacharyya distance and affinity, Mahalanobis distance, and proportional overlap. Choose the metrics you want with metrics; the default computes all of them. Results are returned globally or by group, in a wide format (one column per metric, the default) or a tidy long format (one row per metric per comparison). The percent_overlap values are 0–1 proportions, not 0–100 percentages.

Usage

phontrast(
  data,
  features,
  category_col,
  group_col = NULL,
  metrics = c("jsd", "js_distance", "pillai", "bhattacharyya", "mahalanobis", "overlap"),
  min_tokens = 20,
  bw = c("Hpi", "Hscv", "Hpi.diag", "scott.diag"),
  eval_on = c("pooled", "group1", "group2", "pooled_sample"),
  eval_n = NULL,
  eval_seed = NULL,
  engine = c("ks", "fast_diag", "fast_diagonal"),
  chunk_size = 1000L,
  eps = 1e-06,
  output = c("wide", "long"),
  do_boot = FALSE,
  n_boot = 1000,
  conf_level = 0.95,
  progress = TRUE,
  method = c("mc", "legacy"),
  density = c("kde", "mvnorm"),
  mc_n = 10000L,
  bw_scale = 1
)

Arguments

data

Data frame containing category labels and acoustic features.

features

Character vector of numeric feature columns.

category_col

String; column giving the two categories to compare.

group_col

Optional character vector of one or more grouping columns. If NULL, metrics are computed globally. Multiple grouping columns are combined into a labeled group value such as "Sex=F | Style=read".

metrics

Character vector selecting which contrast metrics to compute. Any of "jsd", "js_distance", "pillai", "bhattacharyya", "mahalanobis", "overlap", "tv", "bhattacharyya_kde", and "euclidean". Defaults to the first six, the metric set of earlier releases. "bhattacharyya" returns both the Bhattacharyya distance and affinity under a closed-form multivariate-normal fit. The remaining three are opt-in: "tv" is total variation, 1 - proportional overlap (total_variation); "bhattacharyya_kde" returns the Bhattacharyya distance and affinity and the Hellinger distance (bhatt_kde_dist, bhatt_kde_affinity, hellinger) read off the same kernel densities as the Jensen-Shannon and overlap columns – a matched-kernel estimator of the quantity the closed-form column estimates parametrically, so measure and estimator can be told apart; and "euclidean" is the Euclidean distance between the two category means after dividing each feature by the standard deviation of the pooled two-category sample (euclidean_dist, bounded by 2\sqrt{d} for equal category sizes). All kernel-family columns (Jensen-Shannon, overlap, total variation, kernel Bhattacharyya, Hellinger) come from one shared density estimate per comparison.

min_tokens

Minimum tokens required globally or per group.

bw

Bandwidth selection method passed to jsd_kde_nd() and percent_overlap_kde().

eval_on

KDE evaluation points passed to jsd_kde_nd() and percent_overlap_kde().

eval_n

Optional maximum number of KDE evaluation points passed to jsd_kde_nd() and percent_overlap_kde().

eval_seed

Optional integer seed for KDE evaluation-point subsampling.

engine

KDE evaluation engine passed to jsd_kde_nd() and percent_overlap_kde(). "fast_diagonal" is accepted as an alias for "fast_diag".

chunk_size

Chunk size for engine = "fast_diag".

eps

Small ridge constant for covariance-based metrics.

output

Output format: "wide" returns one row per global/group comparison; "long" returns one row per metric per comparison.

do_boot

Logical; if TRUE, compute bootstrap means, standard deviations, and confidence intervals for each reported metric.

n_boot

Number of bootstrap resamples if do_boot = TRUE.

conf_level

Confidence level for bootstrap intervals.

progress

Logical; if TRUE, print progress messages while bootstrap resamples are running.

method

KDE estimator for the JSD and percent-overlap columns, passed to jsd_kde_nd()/percent_overlap_kde(): "mc" (default) for the Monte-Carlo plug-in, or "legacy" for the pre-1.2.0 self-normalized estimate. Ignored when density = "mvnorm".

density

Density model behind the two distributional metrics (Jensen-Shannon and proportional overlap): "kde" (default) estimates each category's density by kernel density estimation; "mvnorm" fits one multivariate normal per category and estimates those two metrics between the fitted Gaussians by Monte-Carlo. This lets the density estimator be matched to the same multivariate-normal assumptions the Pillai, Bhattacharyya, and Mahalanobis columns already make. The Pillai, Bhattacharyya, and Mahalanobis columns are parametric by construction and are unaffected by this argument.

mc_n

Positive integer; number of Monte-Carlo samples drawn from each fitted Gaussian for the Jensen-Shannon and overlap columns when density = "mvnorm" (default 10000). Ignored when density = "kde".

bw_scale

Positive number multiplying the selected kernel bandwidth on the standard-deviation scale for the Jensen-Shannon and overlap columns (default 1); 0.5 and 2 give the halved and doubled bandwidths of the smoothing-sensitivity check in rank_contrasts(). Ignored when density = "mvnorm".

Details

Use estimate_jsd() when Jensen-Shannon divergence is the only outcome of interest, and the lower-level metric helpers when you need direct control over one estimator.

Metric directions differ. JSD, Jensen-Shannon distance, Pillai trace, Bhattacharyya distance, and Mahalanobis distance increase as categories become more separated. Percent overlap and Bhattacharyya affinity increase as categories overlap more. Long output includes orientation, separation_value, and separation_rank columns so all metrics can be read on a separation-oriented scale.

If do_boot = TRUE, each metric is recomputed on n_boot nonparametric bootstrap resamples to estimate uncertainty. This can take substantial time because every resample recomputes KDE, MANOVA, and covariance-based metrics. Progress messages are printed by default while bootstrapping is running; set progress = FALSE to suppress them.

Value

A data frame containing only the requested metrics. Wide output (the default) contains one column per requested metric plus pillai_p_value when Pillai is requested; with do_boot = TRUE it also includes metric-specific *_mean, *_sd, *_ci_lower, *_ci_upper, and *_n_boot columns. Long output contains metric, estimate, orientation, bounded_0_1, separation_value, separation_rank, and p_value (populated for the Pillai row, NA otherwise) columns; with do_boot = TRUE it also includes boot_mean, boot_sd, ci_lower, ci_upper, n_boot, and conf_level. The result carries class "phontrast_contrast", so plot() and ggplot2::autoplot() draw it directly via plot_overlap_metrics().

Examples

set.seed(2026)
vowels <- data.frame(
  speaker = rep(c("s01", "s02"), each = 60),
  vowel = rep(rep(c("ih", "eh"), each = 30), 2),
  f1 = c(
    rnorm(30, 500, 55), rnorm(30, 560, 60),
    rnorm(30, 510, 60), rnorm(30, 575, 65)
  ),
  f2 = c(
    rnorm(30, 1980, 150), rnorm(30, 1880, 155),
    rnorm(30, 1960, 160), rnorm(30, 1840, 165)
  )
)

# All metrics in one wide comparison table (the default), by speaker.
phontrast(
  data = vowels,
  features = c("f1", "f2"),
  category_col = "vowel",
  group_col = "speaker"
)

# A single metric in wide format.
phontrast(
  data = vowels,
  features = c("f1", "f2"),
  category_col = "vowel",
  group_col = "speaker",
  metrics = "pillai",
  output = "wide"
)

# Bootstrapping is useful but slower because every requested metric is
# recomputed on every resample. Use a larger n_boot for real analyses.
phontrast(
  data = vowels,
  features = "f1",
  category_col = "vowel",
  group_col = "speaker",
  metrics = c("jsd", "pillai"),
  do_boot = TRUE,
  n_boot = 5,
  progress = FALSE
)

phontrast documentation built on Oct. 7, 2026, 5:06 p.m.