rank_contrasts: Rank speakers' contrasts by Jensen-Shannon distance and check...

View source: R/rank_contrasts.R

rank_contrastsR Documentation

Rank speakers' contrasts by Jensen-Shannon distance and check them against Pillai

Description

One-call implementation of the measurement protocol recommended by the simulation study behind phontrast (Berry, under review, Sec. VII.A). For each speaker (or other unit named by group_col) it

  1. computes Jensen-Shannon distance (\sqrt{JSD}) and the Pillai trace on the same tokens, from one shared kernel density estimate, and reports the estimated shared probability mass beside them;

  2. ranks the speakers by \sqrt{JSD} on the percentile-rank scale of percentile_rank();

  3. flags speakers whose Pillai percentile rank differs from their \sqrt{JSD} percentile rank by margin (0.25 of the ordering) or more, for inspection with inspect_contrast() or plot_contrast().

The conditions the study attaches to those steps are applied and reported rather than left to the user: measurements at the \sqrt{JSD} ceiling are set apart, sample-size floors decide whether a rank or a flag may be read, and a bandwidth check marks measurements whose rank depends on the smoothing.

Usage

rank_contrasts(
  data,
  features,
  category_col,
  group_col,
  margin = 0.25,
  ceiling = 0.99,
  bw_check = TRUE,
  estimator = recommended_estimator(length(features)),
  min_tokens = 10,
  eval_seed = NULL,
  chunk_size = 1000L
)

Arguments

data

Data frame with the speaker, category, and feature columns.

features

Character vector of numeric feature columns.

category_col

String; column giving the two categories of the contrast (for example the two vowels).

group_col

Character vector of one or more columns identifying the speakers (or other units) to rank. Required: one speaker measured once on one contrast gives nothing to rank.

margin

Inspection margin on the percentile-rank scale: the flag is raised when abs(rank_diff) >= margin and a measurement is set aside by the bandwidth check when its \sqrt{JSD} rank moves by margin or more. Default 0.25 of the ordering, the value calibrated in the study.

ceiling

\sqrt{JSD} at or above which a measurement counts as at the ceiling (default 0.99).

bw_check

Logical; run the bandwidth check (default TRUE).

estimator

Kernel estimator settings: a list with elements bw, engine, eval_n, and loo as returned by recommended_estimator(). Defaults to the settings the study used at this dimensionality, recommended_estimator(length(features)).

min_tokens

Minimum tokens in the smaller category for a speaker to be measured at all (default 10). Speakers below it, or without exactly two observed categories, are left out with a message; the licensing floors above are applied to the speakers that remain.

eval_seed

Optional integer seed used when estimator$eval_n subsamples the evaluation points; makes the estimates reproducible.

chunk_size

Chunk size for engine = "fast_diag".

Value

A tibble of class "phontrast_ranking", one row per measured speaker, sorted by sqrt_jsd (largest first), with columns group, n_tokens, n_min, sqrt_jsd, pillai, shared_mass, at_ceiling, pr_jsd, pr_pillai, rank_diff (pr_jsd - pr_pillai), flag, rank_licensed, flag_licensed, rank_basis, and, with bw_check = TRUE, sqrt_jsd_half, sqrt_jsd_double, bw_shift, sign_change, and set_aside. The protocol settings, the floors, and the cleaned tokens the ranking was computed from are attached as attr(x, "protocol") (so inspect_contrast() can redraw a speaker without the original data); print() summarizes them. plot() and ggplot2::autoplot() draw the rank-agreement plot.

Ceiling

A measurement with \sqrt{JSD} \ge ceiling (default 0.99) is at the ceiling: the kernel estimate can no longer separate degrees of separation. Its shared mass and Pillai value are reported, but it is excluded from both percentile rankings (at_ceiling = TRUE, pr_jsd and pr_pillai are NA), and the ranks of the remaining k speakers are computed over those k alone.

Sample-size floors

Floors are read from the smaller category's token count per speaker (n_min) and from the number of features d (see protocol_floors()). Ranking by \sqrt{JSD} is licensed from 50 tokens per category at d = 2, 200 at d = 3 or 4, and 500 at d = 5 to 8; above eight dimensions the study gives ordering evidence only, so no rank is licensed. Below the floor rank_basis is "pillai": order that speaker by Pillai. The flag has floors of its own: 100 tokens per category at d = 2 and 200 at d = 3 or 4; from eight dimensions up the agreement test returns agreement whatever the data contain, and it is undefined when d >= 2 * n_min. Where the flag may not be read, flag is NA and rank_diff is still reported. The floors are where the simulation recovers the average ordering of its separation levels; they do not guarantee an accurate ranking of an individual speaker. Eight speakers is the smallest useful set (the calibration used sets of 32); fewer draws a warning.

Bandwidth check

With bw_check = TRUE (the default) \sqrt{JSD} is recomputed at half and at twice the diagonal Scott bandwidth (bw = "scott.diag", bw_scale = 0.5 and 2), at two dimensions as well, where the reported estimate uses the plug-in rule. The two re-estimates are ranked over the same speakers and bw_shift is the percentile-rank change between them. A measurement is set aside (set_aside = TRUE) when abs(bw_shift) >= margin, or when a flagged Pillai–\sqrt{JSD} rank difference changes sign between the halved and the doubled bandwidth (sign_change). The check tests sensitivity to smoothing, not the estimate's accuracy.

See Also

percentile_rank(), protocol_floors(), recommended_estimator(), inspect_contrast(), plot_contrast(), phontrast().

Examples

set.seed(2026)
gaps <- c(30, 60, 90, 120, 150, 180, 210, 240)
cohort <- do.call(rbind, lapply(seq_along(gaps), function(i) {
  data.frame(
    speaker = sprintf("s%02d", i),
    vowel = rep(c("ih", "eh"), each = 50),
    f1 = c(rnorm(50, 500, 55), rnorm(50, 500 + gaps[i], 55)),
    f2 = c(rnorm(50, 1900, 120), rnorm(50, 1900 - gaps[i], 120))
  )
}))
ranking <- rank_contrasts(cohort, c("f1", "f2"), "vowel", "speaker",
                          bw_check = FALSE)
ranking

phontrast documentation built on Oct. 7, 2026, 5:06 p.m.