| pillai_overlap | R Documentation |
Computes the Pillai-Bartlett trace from a MANOVA of features ~ category. Optionally adds separation estimates and a proportion-standardized Pillai score for an exactly two-category, no-covariate design.
pillai_overlap(data, features, category_col, proportion_standardized = FALSE)
data |
Data frame. |
features |
Character vector of numeric feature columns. |
category_col |
String; column giving exactly two categories. |
proportion_standardized |
Logical; if |
With proportion_standardized = FALSE, the function returns the ordinary
Pillai trace and its p-value exactly as before. Raw Pillai describes the
categories in the realized data; the proportion-standardized score targets
the balanced-design score for the same estimated underlying separation.
For two categories with realized counts n_1 and n_2, total
N, p features, error degrees of freedom \nu_e=N-2, and
H=2n_1n_2/N, the optional estimator chain is
\hat D^2 = 2\nu_e[V/(1-V)]/H,
\tilde D^2 = [(\nu_e-p-1)/\nu_e]\hat D^2 - 2p/H,
followed, when \tilde D^2 >= 0, by
V_{eq}=\tilde D^2/(4+\tilde D^2).
The unbiased estimator is as given by Lachenbruch and Mickey (1968), and
V_{eq} implements Becker's (1986) correction and his two-group
multivariate generalization.
If \tilde D^2 < 0, pillai_eq is NA,
pillai_eq_fallback is TRUE, and d2_fallback contains only the
multiplicatively corrected first term. This common near-merger outcome is
labelled separately because the fallback retains the split-dependent bias
bias_2p_over_H and is not a balanced-design equivalent.
The estimator chain assumes multivariate normality within each category and
a common within-category covariance. It supports no covariates. Both
categories must contain at least two complete tokens, and the within-class
error SSCP must be nonsingular. When \nu_e-p-1 <= 0, all optional
fields are returned as typed NA values with a warning. A minority category
with fewer than p+1 tokens remains computable but is flagged by
fragile_minority.
Because x/(4+x) is strictly concave, pillai_eq is slightly downward
biased for the balanced-design target. This bias grows with the sampling
variance of \tilde D^2 and is largest for small, imbalanced,
near-merged samples.
A list with elements pillai and p_value. When
proportion_standardized = TRUE, the list additionally contains n1,
n2, H, d2_plugin, d2_unbiased, pillai_eq,
pillai_eq_fallback, d2_fallback, bias_2p_over_H, and
fragile_minority. Counts follow the category factor-level order used by
the model. On the fallback path, pillai_eq is NA and
d2_fallback is labelled separately; off that path, d2_fallback is
NA.
Becker, G. (1986). Correcting the point-biserial correlation for attenuation owing to unequal sample size. Journal of Experimental Education, 55(1), 5-8.
Lachenbruch, P. A., & Mickey, M. R. (1968). Estimation of error rates in discriminant analysis. Technometrics, 10(1), 1-11.
Berry, G. M. (2026). Beyond the null: Calibration, balance, and the interpretation of Pillai scores. Preprint.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.