| pseudobulk_deseq2 | R Documentation |
Runs DESeq2::DESeq() on a feature-by-pseudobulk count matrix to
identify genes whose expression differs between groups of
pseudobulks. Three test designs are supported, selected by mode:
"one_vs_all" (default) - test each level of the contrast
variable against the union of all others. Useful for marker
discovery. Dispatched to pseudobulk_one_vs_all().
"pairwise" - test every ordered pair of levels separately.
Useful for high-confidence markers; combine with
summarize_dge_pairs() to keep the most conservative pair.
Dispatched to pseudobulk_pairwise().
"within" - split on the first term of dge_formula and
test the second term within each split level. Useful for
condition / case-vs-control comparisons restricted to one
cluster at a time. Dispatched to pseudobulk_within().
Pseudobulk inputs are typically produced by collapse_counts().
Genes with low counts across pseudobulks are filtered out before
fitting (controlled by min_counts_per_sample and
present_in_min_samples). See the pseudobulk vignette for an
end-to-end walkthrough on a real single-cell dataset.
pseudobulk_deseq2(
dge_formula,
meta_data,
counts_df,
verbose = TRUE,
min_counts_per_sample = 10,
present_in_min_samples = 5,
collapse_background = TRUE,
vals_test = NULL,
mode = c("one_vs_all", "pairwise", "within")[1]
)
dge_formula |
One-sided formula such as |
meta_data |
data.frame of pseudobulk metadata. One row per
pseudobulk; should contain only the variables used in
|
counts_df |
Feature-by-pseudobulk integer count matrix. Rows
are features; columns must align with rows of |
verbose |
Logical. Print progress messages. Default |
min_counts_per_sample |
Minimum count per pseudobulk for a
gene to be considered expressed in that pseudobulk. Default |
present_in_min_samples |
Minimum number of pseudobulks in
which a gene must reach |
collapse_background |
Used only when |
vals_test |
Character vector of contrast levels to test. If
|
mode |
One of |
A long-form data.frame of DESeq2 results. Columns include
the group identifier(s) (group, or group1 / group2 in
pairwise mode), feature, and the DESeq2 columns baseMean,
log2FoldChange, lfcSE, stat, pvalue, padj.
collapse_counts(), top_markers_dds(),
summarize_dge_pairs()
if (requireNamespace("DESeq2", quietly = TRUE)) {
## 40 genes x 300 cells from 2 clusters across 6 donors
m <- matrix(sample.int(8, 40 * 300, replace = TRUE), nrow = 40)
rownames(m) <- paste0("G", 1:40)
colnames(m) <- paste0("C", 1:300)
meta <- data.frame(
cluster = sample(c("a", "b"), 300, replace = TRUE),
donor = sample(paste0("d", 1:6), 300, replace = TRUE)
)
## collapse cells into per-(cluster, donor) pseudobulks
dc <- collapse_counts(m, meta, c("cluster", "donor"))
## test each cluster against the rest
res <- pseudobulk_deseq2(
~cluster,
dc$meta_data["cluster"],
dc$counts_mat,
verbose = FALSE,
present_in_min_samples = 1,
collapse_background = FALSE,
mode = "one_vs_all"
)
head(res)
}
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.