View source: R/compare_overlap_metrics.R
| phontrast | R Documentation |
phontrast() is the package's main entry point. It computes one or more
category separation and overlap metrics for a two-category phonological
contrast in a single call: Jensen-Shannon divergence and distance, the
Pillai-Bartlett trace, Bhattacharyya distance and affinity, Mahalanobis
distance, and proportional overlap. Choose the metrics you want with
metrics; the default computes all of them. Results are returned
globally or by group, in a wide format (one column per metric, the default)
or a tidy long format (one row per metric per comparison). The
percent_overlap values are 0–1 proportions, not 0–100 percentages.
phontrast(
data,
features,
category_col,
group_col = NULL,
metrics = c("jsd", "js_distance", "pillai", "bhattacharyya", "mahalanobis", "overlap"),
min_tokens = 20,
bw = c("Hpi", "Hscv", "Hpi.diag", "scott.diag"),
eval_on = c("pooled", "group1", "group2", "pooled_sample"),
eval_n = NULL,
eval_seed = NULL,
engine = c("ks", "fast_diag", "fast_diagonal"),
chunk_size = 1000L,
eps = 1e-06,
output = c("wide", "long"),
do_boot = FALSE,
n_boot = 1000,
conf_level = 0.95,
progress = TRUE,
method = c("mc", "legacy"),
density = c("kde", "mvnorm"),
mc_n = 10000L,
bw_scale = 1
)
data |
Data frame containing category labels and acoustic features. |
features |
Character vector of numeric feature columns. |
category_col |
String; column giving the two categories to compare. |
group_col |
Optional character vector of one or more grouping columns.
If |
metrics |
Character vector selecting which contrast metrics to compute.
Any of |
min_tokens |
Minimum tokens required globally or per group. |
bw |
Bandwidth selection method passed to |
eval_on |
KDE evaluation points passed to |
eval_n |
Optional maximum number of KDE evaluation points passed to
|
eval_seed |
Optional integer seed for KDE evaluation-point subsampling. |
engine |
KDE evaluation engine passed to |
chunk_size |
Chunk size for |
eps |
Small ridge constant for covariance-based metrics. |
output |
Output format: |
do_boot |
Logical; if |
n_boot |
Number of bootstrap resamples if |
conf_level |
Confidence level for bootstrap intervals. |
progress |
Logical; if |
method |
KDE estimator for the JSD and percent-overlap columns, passed
to |
density |
Density model behind the two distributional metrics
(Jensen-Shannon and proportional overlap): |
mc_n |
Positive integer; number of Monte-Carlo samples drawn from each
fitted Gaussian for the Jensen-Shannon and overlap columns when
|
bw_scale |
Positive number multiplying the selected kernel bandwidth on
the standard-deviation scale for the Jensen-Shannon and overlap columns
(default |
Use estimate_jsd() when Jensen-Shannon divergence is the only outcome
of interest, and the lower-level metric helpers when you need direct control
over one estimator.
Metric directions differ. JSD, Jensen-Shannon distance, Pillai trace,
Bhattacharyya distance, and Mahalanobis distance increase as categories
become more separated. Percent overlap and Bhattacharyya affinity increase
as categories overlap more. Long output includes orientation,
separation_value, and separation_rank columns so all metrics
can be read on a separation-oriented scale.
If do_boot = TRUE, each metric is recomputed on n_boot
nonparametric bootstrap resamples to estimate uncertainty. This can take
substantial time because every resample recomputes KDE, MANOVA, and
covariance-based metrics. Progress messages are printed by default while
bootstrapping is running; set progress = FALSE to suppress them.
A data frame containing only the requested metrics. Wide
output (the default) contains one column per requested metric plus
pillai_p_value when Pillai is requested; with do_boot = TRUE
it also includes metric-specific *_mean, *_sd,
*_ci_lower, *_ci_upper, and *_n_boot columns. Long
output contains metric, estimate, orientation,
bounded_0_1, separation_value, separation_rank, and
p_value (populated for the Pillai row, NA otherwise) columns;
with do_boot = TRUE it also includes boot_mean,
boot_sd, ci_lower, ci_upper, n_boot, and
conf_level. The result carries class "phontrast_contrast",
so plot() and ggplot2::autoplot() draw it directly via
plot_overlap_metrics().
set.seed(2026)
vowels <- data.frame(
speaker = rep(c("s01", "s02"), each = 60),
vowel = rep(rep(c("ih", "eh"), each = 30), 2),
f1 = c(
rnorm(30, 500, 55), rnorm(30, 560, 60),
rnorm(30, 510, 60), rnorm(30, 575, 65)
),
f2 = c(
rnorm(30, 1980, 150), rnorm(30, 1880, 155),
rnorm(30, 1960, 160), rnorm(30, 1840, 165)
)
)
# All metrics in one wide comparison table (the default), by speaker.
phontrast(
data = vowels,
features = c("f1", "f2"),
category_col = "vowel",
group_col = "speaker"
)
# A single metric in wide format.
phontrast(
data = vowels,
features = c("f1", "f2"),
category_col = "vowel",
group_col = "speaker",
metrics = "pillai",
output = "wide"
)
# Bootstrapping is useful but slower because every requested metric is
# recomputed on every resample. Use a larger n_boot for real analyses.
phontrast(
data = vowels,
features = "f1",
category_col = "vowel",
group_col = "speaker",
metrics = c("jsd", "pillai"),
do_boot = TRUE,
n_boot = 5,
progress = FALSE
)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.