wilcoxauc: Fast Wilcoxon rank-sum test and auROC across groups

View source: R/wilcoxauc.R

wilcoxaucR Documentation

Fast Wilcoxon rank-sum test and auROC across groups

Description

For every (feature, group) pair, computes the Wilcoxon rank-sum statistic comparing observations in that group against all other observations, and the area under the ROC curve as a measure of separability. P-values come from the standard Gaussian approximation to the U statistic with a tie correction. Returns one row per (feature, group) with effect-size and percent-expressed columns alongside the test statistics.

Usage

wilcoxauc(X, ...)

## S3 method for class 'seurat'
wilcoxauc(X, ...)

## S3 method for class 'Seurat'
wilcoxauc(
  X,
  group_by = NULL,
  assay = "data",
  groups_use = NULL,
  seurat_assay = "RNA",
  ...
)

## S3 method for class 'SingleCellExperiment'
wilcoxauc(X, group_by = NULL, assay = NULL, groups_use = NULL, ...)

## Default S3 method:
wilcoxauc(
  X,
  y,
  groups_use = NULL,
  verbose = TRUE,
  nthreads = 1,
  transposed = FALSE,
  ...
)

Arguments

X

Input data. One of:

  • a numeric feature-by-observation matrix or data.frame,

  • a sparse dgCMatrix of the same shape,

  • a disk-backed DelayedMatrix (e.g. HDF5-backed, from the DelayedArray / HDF5Array packages), which is processed in feature blocks so the whole matrix never has to be loaded in memory,

  • any other matrix-like class with an as(., "dgCMatrix") coercion method (e.g. BPCells), which is converted up front,

  • a Seurat (v3+) object,

  • a SingleCellExperiment object.

X must not contain NA values. Unlike stats::wilcox.test(), which drops missing values per observation, wilcoxauc() errors on NA input rather than returning silently incorrect results; remove or impute missing values first.

...

Passed to the input-specific method.

group_by

For Seurat and SingleCellExperiment input, name of the metadata column that holds the group labels (e.g. "cluster"). For Seurat, defaults to Idents(X).

assay

For Seurat, the layer name within the selected assay (e.g. "data", "counts", "scale.data"). For SingleCellExperiment, the assay name (e.g. "logcounts", "counts"). Defaults pick a sensible value per input class.

groups_use

Optional character vector restricting the test to a subset of groups in y (or group_by). Default NULL tests every group.

seurat_assay

For Seurat input, the name of the assay to pull from (e.g. "RNA"). Default "RNA".

y

For matrix input, a character/factor vector of group labels with length equal to ncol(X). Ignored for Seurat / SingleCellExperiment input (use group_by instead).

verbose

Logical. Print warnings and informational messages. Default TRUE.

nthreads

Number of threads for the per-feature ranking of sparse (dgCMatrix) input. Default 1 (serial). Values above 1 split the ranking across threads; the result is identical regardless of the thread count. Only sparse input is parallelized – dense matrix and data.frame input are always processed serially. When running under ⁠R CMD check⁠ or on CRAN, keep this at the default so no more than two cores are used.

transposed

Set to TRUE when your observations (cells, samples) are in the rows of X and the features in the columns – i.e. X is the transpose of the default features-by-observations layout. The test then runs directly on that layout without materializing a transposed copy, which saves time and memory on large matrices. Same convention as the transposed argument of scater's calculatePCA() and calculateUMAP(). Only applies to matrix-like input (the Seurat / SingleCellExperiment dispatchers always extract features-by-observations). Default FALSE.

Details

Designed to be fast enough to run on whole-genome × hundred-thousand- cell single-cell matrices in seconds. Sparse dgCMatrix inputs are processed without densification. Convenience dispatchers extract the counts matrix and group labels from Seurat and SingleCellExperiment objects. See the getting-started vignette for an end-to-end example on a real dataset.

Value

table with the following columns:

  • feature - feature name (e.g. gene name).

  • group - group name.

  • avgExpr - mean value of feature in group.

  • logFC - difference of mean feature values between observations in the group vs out of the group. When the input is log-transformed expression (e.g. Seurat's "data" layer or logcounts), this difference of means is a log fold change. On raw (untransformed) values it is a plain difference of means, not a fold change.

  • statistic - Wilcoxon rank sum U statistic.

  • auc - area under the receiver operator curve.

  • pval - nominal p value.

  • padj - Benjamini-Hochberg adjusted p value.

  • pct_in - Percent of observations in the group with non-zero feature value.

  • pct_out - Percent of observations out of the group with non-zero feature value.

See Also

top_markers() to summarize markers per group; pseudobulk_deseq2() for a count-based pseudobulk alternative.

Examples

## generate a tiny toy dataset
set.seed(42)
exprs <- matrix(rpois(25 * 150, lambda = 2), nrow = 25,
                dimnames = list(paste0("G", 1:25), NULL))
y <- rep(c("A", "B", "C"), each = 50)

## on a dense matrix
head(wilcoxauc(exprs, y))

## restrict the comparison to a subset of groups
head(wilcoxauc(exprs, y, c('A', 'B')))

## on a sparse matrix
exprs_sparse <- as(exprs, 'dgCMatrix')
head(wilcoxauc(exprs_sparse, y))

## on a Seurat object (>= v3)
if (requireNamespace("Seurat", quietly = TRUE) &&
    packageVersion("Seurat") >= "3.0") {
    object_seurat <- toy_seurat()
    head(wilcoxauc(object_seurat, 'cell_type'))
}

## on a SingleCellExperiment object
if (requireNamespace("SingleCellExperiment", quietly = TRUE)) {
    object_sce <- toy_sce()
    head(wilcoxauc(object_sce, 'cell_type'))
}


presto documentation built on Sept. 30, 2026, 5:13 p.m.