pillai_overlap: Pillai trace for multivariate overlap

View source: R/pillai_bhatt.R

pillai_overlapR Documentation

Pillai trace for multivariate overlap

Description

Computes the Pillai-Bartlett trace from a MANOVA of features ~ category. Optionally adds separation estimates and a proportion-standardized Pillai score for an exactly two-category, no-covariate design.

Usage

pillai_overlap(data, features, category_col, proportion_standardized = FALSE)

Arguments

data

Data frame.

features

Character vector of numeric feature columns.

category_col

String; column giving exactly two categories.

proportion_standardized

Logical; if TRUE, append the plug-in and unbiased squared Mahalanobis separation estimates and the proportion-standardized Pillai score. The default, FALSE, preserves the original two-field return value.

Details

With proportion_standardized = FALSE, the function returns the ordinary Pillai trace and its p-value exactly as before. Raw Pillai describes the categories in the realized data; the proportion-standardized score targets the balanced-design score for the same estimated underlying separation.

For two categories with realized counts n_1 and n_2, total N, p features, error degrees of freedom \nu_e=N-2, and H=2n_1n_2/N, the optional estimator chain is

\hat D^2 = 2\nu_e[V/(1-V)]/H,

\tilde D^2 = [(\nu_e-p-1)/\nu_e]\hat D^2 - 2p/H,

followed, when \tilde D^2 >= 0, by

V_{eq}=\tilde D^2/(4+\tilde D^2).

The unbiased estimator is as given by Lachenbruch and Mickey (1968), and V_{eq} implements Becker's (1986) correction and his two-group multivariate generalization.

If \tilde D^2 < 0, pillai_eq is NA, pillai_eq_fallback is TRUE, and d2_fallback contains only the multiplicatively corrected first term. This common near-merger outcome is labelled separately because the fallback retains the split-dependent bias bias_2p_over_H and is not a balanced-design equivalent.

The estimator chain assumes multivariate normality within each category and a common within-category covariance. It supports no covariates. Both categories must contain at least two complete tokens, and the within-class error SSCP must be nonsingular. When \nu_e-p-1 <= 0, all optional fields are returned as typed NA values with a warning. A minority category with fewer than p+1 tokens remains computable but is flagged by fragile_minority.

Because x/(4+x) is strictly concave, pillai_eq is slightly downward biased for the balanced-design target. This bias grows with the sampling variance of \tilde D^2 and is largest for small, imbalanced, near-merged samples.

Value

A list with elements pillai and p_value. When proportion_standardized = TRUE, the list additionally contains n1, n2, H, d2_plugin, d2_unbiased, pillai_eq, pillai_eq_fallback, d2_fallback, bias_2p_over_H, and fragile_minority. Counts follow the category factor-level order used by the model. On the fallback path, pillai_eq is NA and d2_fallback is labelled separately; off that path, d2_fallback is NA.

References

Becker, G. (1986). Correcting the point-biserial correlation for attenuation owing to unequal sample size. Journal of Experimental Education, 55(1), 5-8.

Lachenbruch, P. A., & Mickey, M. R. (1968). Estimation of error rates in discriminant analysis. Technometrics, 10(1), 1-11.

Berry, G. M. (2026). Beyond the null: Calibration, balance, and the interpretation of Pillai scores. Preprint.


phontrast documentation built on Oct. 7, 2026, 5:06 p.m.