View source: R/rank_contrasts.R
| rank_contrasts | R Documentation |
One-call implementation of the measurement protocol recommended by the
simulation study behind phontrast (Berry, under review, Sec. VII.A). For
each speaker (or other unit named by group_col) it
computes Jensen-Shannon distance (\sqrt{JSD}) and the Pillai
trace on the same tokens, from one shared kernel density estimate, and
reports the estimated shared probability mass beside them;
ranks the speakers by \sqrt{JSD} on the percentile-rank scale
of percentile_rank();
flags speakers whose Pillai percentile rank differs from their
\sqrt{JSD} percentile rank by margin (0.25 of the ordering)
or more, for inspection with inspect_contrast() or
plot_contrast().
The conditions the study attaches to those steps are applied and reported
rather than left to the user: measurements at the \sqrt{JSD} ceiling
are set apart, sample-size floors decide whether a rank or a flag may be
read, and a bandwidth check marks measurements whose rank depends on the
smoothing.
rank_contrasts(
data,
features,
category_col,
group_col,
margin = 0.25,
ceiling = 0.99,
bw_check = TRUE,
estimator = recommended_estimator(length(features)),
min_tokens = 10,
eval_seed = NULL,
chunk_size = 1000L
)
data |
Data frame with the speaker, category, and feature columns. |
features |
Character vector of numeric feature columns. |
category_col |
String; column giving the two categories of the contrast (for example the two vowels). |
group_col |
Character vector of one or more columns identifying the speakers (or other units) to rank. Required: one speaker measured once on one contrast gives nothing to rank. |
margin |
Inspection margin on the percentile-rank scale: the flag is
raised when |
ceiling |
|
bw_check |
Logical; run the bandwidth check (default |
estimator |
Kernel estimator settings: a list with elements |
min_tokens |
Minimum tokens in the smaller category for a speaker to be measured at all (default 10). Speakers below it, or without exactly two observed categories, are left out with a message; the licensing floors above are applied to the speakers that remain. |
eval_seed |
Optional integer seed used when |
chunk_size |
Chunk size for |
A tibble of class "phontrast_ranking", one row per measured
speaker, sorted by sqrt_jsd (largest first), with columns
group, n_tokens, n_min, sqrt_jsd,
pillai, shared_mass, at_ceiling, pr_jsd,
pr_pillai, rank_diff (pr_jsd - pr_pillai),
flag, rank_licensed, flag_licensed,
rank_basis, and, with bw_check = TRUE,
sqrt_jsd_half, sqrt_jsd_double, bw_shift,
sign_change, and set_aside. The protocol settings, the
floors, and the cleaned tokens the ranking was computed from are attached
as attr(x, "protocol") (so inspect_contrast() can redraw a
speaker without the original data); print()
summarizes them. plot() and ggplot2::autoplot() draw the
rank-agreement plot.
A measurement with \sqrt{JSD} \ge ceiling (default 0.99) is at
the ceiling: the kernel estimate can no longer separate degrees of
separation. Its shared mass and Pillai value are reported, but it is
excluded from both percentile rankings (at_ceiling = TRUE,
pr_jsd and pr_pillai are NA), and the ranks of the
remaining k speakers are computed over those k alone.
Floors are read from the smaller category's token count per speaker
(n_min) and from the number of features d
(see protocol_floors()). Ranking by \sqrt{JSD} is licensed from
50 tokens per category at d = 2, 200 at d = 3 or 4,
and 500 at d = 5 to 8; above eight dimensions the study gives
ordering evidence only, so no rank is licensed. Below the floor
rank_basis is "pillai": order that speaker by Pillai. The
flag has floors of its own: 100 tokens per category at d = 2 and
200 at d = 3 or 4; from eight dimensions up the agreement test
returns agreement whatever the data contain, and it is undefined when
d >= 2 * n_min. Where the flag may not be read, flag is
NA and rank_diff is still reported. The floors are where the
simulation recovers the average ordering of its separation levels; they do
not guarantee an accurate ranking of an individual speaker. Eight speakers
is the smallest useful set (the calibration used sets of 32); fewer draws a
warning.
With bw_check = TRUE (the default) \sqrt{JSD} is recomputed at
half and at twice the diagonal Scott bandwidth (bw = "scott.diag",
bw_scale = 0.5 and 2), at two dimensions as well, where the
reported estimate uses the plug-in rule. The two re-estimates are ranked
over the same speakers and bw_shift is the percentile-rank change
between them. A measurement is set aside (set_aside = TRUE) when
abs(bw_shift) >= margin, or when a flagged Pillai–\sqrt{JSD}
rank difference changes sign between the halved and the doubled bandwidth
(sign_change). The check tests sensitivity to smoothing, not the
estimate's accuracy.
percentile_rank(), protocol_floors(),
recommended_estimator(), inspect_contrast(),
plot_contrast(), phontrast().
set.seed(2026)
gaps <- c(30, 60, 90, 120, 150, 180, 210, 240)
cohort <- do.call(rbind, lapply(seq_along(gaps), function(i) {
data.frame(
speaker = sprintf("s%02d", i),
vowel = rep(c("ih", "eh"), each = 50),
f1 = c(rnorm(50, 500, 55), rnorm(50, 500 + gaps[i], 55)),
f2 = c(rnorm(50, 1900, 120), rnorm(50, 1900 - gaps[i], 120))
)
}))
ranking <- rank_contrasts(cohort, c("f1", "f2"), "vowel", "speaker",
bw_check = FALSE)
ranking
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.