| desc_stats | R Documentation |
desc_stats() computes descriptive statistics for numeric columns, overall
or by group, in one tidy table (one row per group x variable, or one row
per group in wide shape). The statistic sets follow
rstatix::get_summary_stats(). An optional total row gives the
statistics of every variable over all records; results can be returned
as numbers or as report-ready text such as "26.66 ± 4.51".
desc_stats(
data,
cols = NULL,
by = NULL,
type = "common",
stats = NULL,
total = FALSE,
total_label = "Total",
min_n = 1L,
probs = c(0, 0.25, 0.5, 0.75, 1),
conf_level = 0.95,
na_rm = TRUE,
digits = NULL,
fmt = NULL,
labels = NULL,
shape = "long",
sep = "_",
out_type = "dt"
)
data |
A |
cols |
Numeric columns to summarise, as names or indices. Default
|
by |
Grouping column(s), as names or indices. Default |
type |
Preset set of statistics (ignored when
|
stats |
Optional character vector of statistics, overriding |
total |
Logical. If
Without |
total_label |
Label of the total row and the total variable. Default
|
min_n |
Minimum number of (non-missing) values a group needs for its
statistics to be reported. Groups with fewer values keep |
probs |
Probabilities for |
conf_level |
Confidence level of |
na_rm |
Logical. If |
digits |
|
fmt |
The statistics needed by the templates are computed automatically;
|
labels |
|
shape |
|
sep |
Separator between variable and statistic in wide column names.
Default |
out_type |
|
Definitions: q1 / q3 and quantile use stats::quantile()
(type 7), iqr = q3 - q1, mad is stats::mad() (scaled by 1.4826),
se = sd / sqrt(n), ci is the half-width of the t-based confidence
interval of the mean (qt((1 + conf_level) / 2, n - 1) * se), and
ci_low / ci_high are its bounds (mean -/+ ci),
cv = sd / mean as a ratio (shown in percent with fmt = "{cv:1\%}"),
skew is the adjusted Fisher-Pearson skewness G1 and kurt the excess
kurtosis G2, as in SAS, SPSS and Excel's SKEW() / KURT() (NA with
fewer than 3 / 4 values).
The CV is only meaningful for strictly positive data and is NA when a
group contains values <= 0. Statistics that are undefined for a group
(e.g. sd with one value) are NA.
Computation: statistics are computed column by column on the input table
(no reshaping to long format, so memory use stays close to the input
size). n, mean, sd, min, max and median use data.table's
optimised grouped functions, quantiles are read off once-sorted groups,
so tens of thousands of groups (sires, litters, pens) are summarised in
about a second per million records.
Rows are ordered by variable (in the order of cols),
then by group (in the original order of the by values: numeric order,
factor levels, alphabetical for text), with the total row last. With a
total row, by columns that are not factors or character are returned as
character; factors gain the label as their last level.
A data.table (or data.frame). Long shape: the by columns,
variable and one column per statistic (or per fmt template, as
text). Wide shape: the by columns followed by one column per
variable x statistic (or template).
top_perc() for statistics of the top / bottom X% per group.
# Example 1: Common statistics for every numeric column of iris
desc_stats(iris)
# Example 2: Mean and SD by group, with a total row over all records
desc_stats(
iris,
by = "Species", # Grouping column
type = "mean_sd", # Preset: n, mean, sd
total = TRUE, # n = sum of the groups, mean / sd of all records
digits = 2 # Round the statistics
)
# Example 3: Count table with row and column totals
# (counts are the only statistic that adds up across variables)
desc_stats(
iris,
by = "Species",
stats = "n",
total = TRUE, # Total row and Total column
shape = "wide"
)
# Without groups the total adds up the counts of all variables
desc_stats(iris, type = "mean_sd", total = TRUE, digits = 2)
# Example 3b: Report table - groups in rows, traits in columns, "mean ± sd"
desc_stats(
mtcars,
cols = c("mpg", "hp", "wt"),
by = "cyl",
fmt = "{mean} ± {sd}", # Report-ready text
total = TRUE,
shape = "wide" # One row per group, one column per trait
)
# Example 4: Several templates, per-placeholder decimals, CV in percent and
# percentiles
desc_stats(
mtcars,
cols = c("mpg", "wt"),
by = "am",
fmt = c(N = "{n}",
"Mean ± SD" = "{mean:1} ± {sd:1}",
"CV" = "{cv:1%}",
"Median [P2.5, P97.5]" = "{median:1} [{q2.5:1}, {q97.5:1}]"),
total = TRUE
)
# Example 5: Quantiles, returned as a data.frame
desc_stats(iris, cols = 1:2, type = "quantile",
probs = c(0.05, 0.5, 0.95), out_type = "df")
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.