phontrast-package: phontrast: Contrast and Separation Metrics for Phonological...

phontrast-packageR Documentation

phontrast: Contrast and Separation Metrics for Phonological Categories

Description

A unified toolkit for quantifying the separation and overlap between phonological categories (e.g., vowels, consonants) in arbitrary n-dimensional acoustic spaces such as formants, MFCCs, spectral features, duration, or learned embeddings. The main entry point, phontrast(), computes and compares multiple contrast metrics in one call: Jensen-Shannon divergence and distance, the Pillai-Bartlett trace, Bhattacharyya distance and affinity, Mahalanobis distance, and proportional overlap, globally or by group, with optional bootstrap intervals. Percent overlap is returned as a 0–1 proportion, not a 0–100 percentage.

Details

phontrast was formerly released as phonJSD (through version 1.2.0); see the package NEWS for migration notes.

Recommended workflow

  1. Start with phontrast() to compute one or more contrast metrics for a two-category contrast, globally or by group, in tidy long or wide form.

  2. Use estimate_jsd() when Jensen-Shannon divergence or Jensen-Shannon distance is the primary outcome and you need optional bootstrap intervals.

  3. Use lower-level helpers such as jsd_kde_nd(), percent_overlap_kde(), pillai_overlap(), and bhattacharyya_mvnorm() when validating methods, debugging one contrast, or reproducing a specific metric.

  4. Use plot_contrast() for a distribution-aware, annotated view of one contrast that draws the same density model the metrics use, and plot_overlap_metrics() (also available as plot() / ggplot2::autoplot() on phontrast() results), plot_category_space(), and plot_category_pca() for ggplot2-backed diagnostics and presentation figures, all sharing the colorblind-safe theme_phontrast() visual identity.

  5. To rank a set of speakers on one contrast, use rank_contrasts() (see the next section), then plot() the ranking and inspect_contrast() any flagged speaker.

Ranking protocol

rank_contrasts() implements the measurement protocol recommended by the simulation study behind this package (Berry, under review, Sec. VII.A; see citation("phontrast")): compute Jensen-Shannon distance and Pillai on the same tokens and report shared mass beside them; rank speakers by Jensen-Shannon distance on the percentile-rank scale of percentile_rank(); and flag speakers whose Pillai and Jensen-Shannon percentile ranks differ by 0.25 of the ordering or more. Measurements at the Jensen-Shannon ceiling are set apart, sample-size floors (protocol_floors()) decide per speaker whether a rank or a flag may be read, and a bandwidth check marks measurements whose rank depends on the smoothing. recommended_estimator() supplies the kernel settings the study used at each dimensionality; plot_rank_agreement() and inspect_contrast() draw the ranking and a flagged speaker; the bundled vowel_cohort data and the vignette "Ranking speakers by Jensen-Shannon distance and checking Pillai agreement" walk through it.

Choosing metrics

JSD, Jensen-Shannon distance, Pillai trace, Bhattacharyya distance, and Mahalanobis distance increase as categories become more separated. Percent overlap and Bhattacharyya affinity increase as categories overlap more. The long output from phontrast() includes an orientation column and a separation-oriented separation_value column to make these directions explicit. JSD and percent overlap estimate distributional separation/overlap using KDE by default; Pillai and Mahalanobis emphasize mean separation; Bhattacharyya metrics use a multivariate-normal approximation. Opt-in metrics add total variation ("tv"), the matched-kernel Bhattacharyya and Hellinger distances read off the same kernel densities as JSD ("bhattacharyya_kde"), and the Euclidean distance between standardized category means ("euclidean").

Density backends

The distributional metrics (Jensen-Shannon divergence and proportional overlap) are computed from a density estimate for each category. The density argument selects that estimate: "kde" (the default) uses kernel density estimation, and "mvnorm" fits one multivariate normal per category and estimates the metric between the two Gaussians by Monte-Carlo (with mc_n samples and reproducible eval_seed). The "mvnorm" backend matches the estimator behind JSD and overlap to the same multivariate-normal assumptions the Pillai, Bhattacharyya, and Mahalanobis columns already make, and is convenient for higher-dimensional feature spaces where multivariate KDE is impractical. It is available on phontrast(), estimate_jsd(), estimate_overlap(), jsd_summary(), global_boot_jsd(), jsd_kde_nd(), and percent_overlap_kde(); the parametric metrics are unaffected by it.

High-dimensional workflows

Metrics can be estimated in arbitrary n-dimensional numeric feature spaces, including MFCCs and learned embeddings. Use plot_category_pca() for a two-dimensional PCA diagnostic, but report metric estimates from the intended full feature set.

Confidence intervals

Confidence intervals use ci_lower and ci_upper columns. Legacy JSD aliases jsd_low and jsd_high are retained for compatibility.

Author(s)

Maintainer: Grant M. Berry berry.grant@gmail.com

See Also

Useful links:


phontrast documentation built on Oct. 7, 2026, 5:06 p.m.