| phontrast-package | R Documentation |
A unified toolkit for quantifying the separation and overlap between
phonological categories (e.g., vowels, consonants) in arbitrary
n-dimensional acoustic spaces such as formants, MFCCs, spectral features,
duration, or learned embeddings. The main entry point, phontrast(),
computes and compares multiple contrast metrics in one call: Jensen-Shannon
divergence and distance, the Pillai-Bartlett trace, Bhattacharyya distance
and affinity, Mahalanobis distance, and proportional overlap, globally or by
group, with optional bootstrap intervals. Percent overlap is returned as a
0–1 proportion, not a 0–100 percentage.
phontrast was formerly released as phonJSD (through version 1.2.0); see the package NEWS for migration notes.
Start with phontrast() to compute one or more contrast metrics
for a two-category contrast, globally or by group, in tidy long or wide
form.
Use estimate_jsd() when Jensen-Shannon divergence or
Jensen-Shannon distance is the primary outcome and you need optional
bootstrap intervals.
Use lower-level helpers such as jsd_kde_nd(),
percent_overlap_kde(), pillai_overlap(), and
bhattacharyya_mvnorm() when validating methods, debugging one
contrast, or reproducing a specific metric.
Use plot_contrast() for a distribution-aware, annotated view
of one contrast that draws the same density model the metrics use, and
plot_overlap_metrics() (also available as plot() /
ggplot2::autoplot() on phontrast() results),
plot_category_space(), and plot_category_pca() for
ggplot2-backed diagnostics and presentation figures, all sharing the
colorblind-safe theme_phontrast() visual identity.
To rank a set of speakers on one contrast, use
rank_contrasts() (see the next section), then plot() the
ranking and inspect_contrast() any flagged speaker.
rank_contrasts() implements the measurement protocol recommended by
the simulation study behind this package (Berry, under review, Sec. VII.A;
see citation("phontrast")): compute Jensen-Shannon distance and
Pillai on the same tokens and report shared mass beside them; rank speakers
by Jensen-Shannon distance on the percentile-rank scale of
percentile_rank(); and flag speakers whose Pillai and Jensen-Shannon
percentile ranks differ by 0.25 of the ordering or more. Measurements at the
Jensen-Shannon ceiling are set apart, sample-size floors
(protocol_floors()) decide per speaker whether a rank or a flag may
be read, and a bandwidth check marks measurements whose rank depends on the
smoothing. recommended_estimator() supplies the kernel settings the
study used at each dimensionality; plot_rank_agreement() and
inspect_contrast() draw the ranking and a flagged speaker; the
bundled vowel_cohort data and the vignette "Ranking speakers by
Jensen-Shannon distance and checking Pillai agreement" walk through it.
JSD, Jensen-Shannon distance, Pillai trace, Bhattacharyya distance, and
Mahalanobis distance increase as categories become more separated. Percent
overlap and Bhattacharyya affinity increase as categories overlap more. The
long output from phontrast() includes an orientation column and
a separation-oriented separation_value column to make these directions
explicit. JSD and percent overlap estimate distributional separation/overlap
using KDE by default; Pillai and Mahalanobis emphasize mean separation;
Bhattacharyya metrics use a multivariate-normal approximation. Opt-in
metrics add total variation ("tv"), the matched-kernel Bhattacharyya
and Hellinger distances read off the same kernel densities as JSD
("bhattacharyya_kde"), and the Euclidean distance between
standardized category means ("euclidean").
The distributional metrics (Jensen-Shannon divergence and proportional
overlap) are computed from a density estimate for each category. The
density argument selects that estimate: "kde" (the default)
uses kernel density estimation, and "mvnorm" fits one multivariate
normal per category and estimates the metric between the two Gaussians by
Monte-Carlo (with mc_n samples and reproducible eval_seed).
The "mvnorm" backend matches the estimator behind JSD and overlap to
the same multivariate-normal assumptions the Pillai, Bhattacharyya, and
Mahalanobis columns already make, and is convenient for higher-dimensional
feature spaces where multivariate KDE is impractical. It is available on
phontrast(), estimate_jsd(), estimate_overlap(),
jsd_summary(), global_boot_jsd(), jsd_kde_nd(), and
percent_overlap_kde(); the parametric metrics are unaffected by it.
Metrics can be estimated in arbitrary n-dimensional numeric feature spaces,
including MFCCs and learned embeddings. Use plot_category_pca() for a
two-dimensional PCA diagnostic, but report metric estimates from the intended
full feature set.
Confidence intervals use ci_lower and ci_upper columns.
Legacy JSD aliases jsd_low and jsd_high are retained for
compatibility.
Maintainer: Grant M. Berry berry.grant@gmail.com
Useful links:
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.