filter.asv: Filter ASV count matrix by library size (samples) and...

View source: R/filter_asv.R

filter.asvR Documentation

Filter ASV count matrix by library size (samples) and prevalence (features)

Description

Filters an ASV count matrix by: (1) removing samples with library size below min.lib; (2) keeping features (taxa) that are "present" in at least ceiling(prev.prop * nrow(S.counts)) samples, where presence is defined by either min.count (counts) or min.rel (relative). After feature filtering, zero-total samples are dropped and a row-normalized relative-abundance matrix is returned alongside the filtered counts.

Usage

filter.asv(
  S.counts,
  min.lib = 1000,
  prev.prop = 0.05,
  min.count = 2,
  min.rel = NULL,
  min.feat.total = NULL,
  verbose = TRUE
)

Arguments

S.counts

Numeric matrix (samples x features) of nonnegative counts.

min.lib

Integer. Minimum library size (row sum of counts) to keep a sample. Default: 1000.

prev.prop

Numeric in (0,1]. Minimum fraction of samples where a feature must be present. Default: 0.05.

min.count

Integer (>=1) or NULL. Reads to call "present" (ignored if min.rel set). Default: 2.

min.rel

Numeric in (0,1) or NULL. Relative abundance to call "present" (overrides min.count). Default: NULL.

min.feat.total

Integer (>=0) or NULL. Optional minimum total reads across all samples per feature. Default: NULL.

verbose

Logical. Print keep/drop summaries. Default: TRUE.

Details

Sample filtering is applied on raw counts first. Prevalence is computed on the post-sample-filter matrix using either a count or relative rule. A feature is retained if prevalence \ge \lceil \text{prev.prop} \times n_{\text{samples}} \rceil. After feature filtering, samples with zero remaining counts are dropped and the relative matrix rel is row-normalized.

Value

A list with elements:

counts

Filtered count matrix (samples x features).

rel

Row-normalized relative-abundance matrix.

kept.sample.idx

Kept sample indices (original order).

kept.feature.idx

Kept feature indices (original order).

prevalence

Per-feature prevalence counts for kept features.

thresholds

List of thresholds actually used.

Conventions

Dot-delimited names; rows are samples, columns are features. Supply raw counts.

Examples

set.seed(1)
S <- matrix(rpois(100 * 20, lambda = 5), nrow = 100, ncol = 20)
res <- filter.asv(S, min.lib = 50, prev.prop = 0.1, min.count = 2)
dim(res$counts); dim(res$rel)
range(rowSums(res$rel))


linf documentation built on Aug. 5, 2026, 9:08 a.m.