vowel_cohort: Simulated twelve-speaker vowel cohort

vowel_cohortR Documentation

Simulated twelve-speaker vowel cohort

Description

A simulated cohort for demonstrating the ranking protocol of rank_contrasts(): twelve speakers, two vowel categories ("ih" and "eh"), and two acoustic features (F1 and F2, in Hz). Each speaker was built to exercise one part of the protocol:

spk01–spk08

Gaussian categories whose centroid gap grows from 25 to 200 Hz, 100 tokens per vowel: a graded ordering that both Jensen-Shannon distance and Pillai recover.

spk09

The same centroids for both vowels, but a bimodal "eh" (two variants 190 Hz apart in F1 and 440 Hz apart in F2), 100 tokens per vowel. A mean-based measure sees no contrast while the distributions barely overlap: the planted Pillai / \sqrt{JSD} disagreement the agreement flag should catch.

spk10

Fully separated categories, 100 tokens per vowel: \sqrt{JSD} at the ceiling.

spk11

60 tokens per vowel: ranking by \sqrt{JSD} is licensed at two dimensions (floor 50) but the flag is not readable (floor 100).

spk12

40 tokens per vowel: below the two-dimensional rank floor, so the speaker is ordered by Pillai.

Usage

vowel_cohort

Format

A data frame with 2200 rows and 4 columns:

speaker

Character; speaker identifier "spk01" to "spk12".

vowel

Character; vowel category, "ih" or "eh".

f1

Numeric; first formant frequency in Hz.

f2

Numeric; second formant frequency in Hz.

Source

Simulated with a fixed seed; the generating script is data-raw/vowel_cohort.R in the package repository.

See Also

rank_contrasts(), inspect_contrast().

Examples

head(vowel_cohort)
table(vowel_cohort$speaker, vowel_cohort$vowel)

phontrast documentation built on Oct. 7, 2026, 5:06 p.m.