tab: Create Descriptive, Comparative, and Regression Tables

View source: R/zzz-r4vn-tab-hierarchical.R View source: R/tab.R

tabR Documentation

Create Descriptive, Comparative, and Regression Tables

Description

Creates publication-style tables for descriptive analysis, group comparisons, binary-outcome regression, and continuous-outcome linear regression.

Usage

tab(
  ...,
  data = NULL,
  vars = NULL,
  by = NULL,
  superby = NULL,
  digit = 1,
  p_digit = 3,
  effect_digit = 2,
  missing = "ifany",
  row = FALSE,
  col = TRUE,
  cell = FALSE,
  overall = "first",
  descriptive = TRUE,
  rvrow = NULL,
  rvcol = FALSE,
  test = TRUE,
  pvalue = TRUE,
  bold_p = TRUE,
  p_bold = 0.05,
  test_note = TRUE,
  interaction = TRUE,
  or = FALSE,
  rr = FALSE,
  pr = FALSE,
  event = NULL,
  adjusted = NULL,
  multi = NULL,
  effect_ref = NULL,
  template = c("journal", "clean", "minimal"),
  append = NULL,
  file = NULL,
  raw = FALSE,
  name = FALSE,
  title = NULL,
  show = TRUE,
  mode = c("auto", "console", "table")
)

tab(
  ...,
  data = NULL,
  vars = NULL,
  by = NULL,
  superby = NULL,
  digit = 1,
  p_digit = 3,
  effect_digit = 2,
  missing = "ifany",
  row = FALSE,
  col = TRUE,
  cell = FALSE,
  overall = "first",
  descriptive = TRUE,
  rvrow = NULL,
  rvcol = FALSE,
  test = TRUE,
  pvalue = TRUE,
  bold_p = TRUE,
  p_bold = 0.05,
  test_note = TRUE,
  interaction = TRUE,
  or = FALSE,
  rr = FALSE,
  pr = FALSE,
  event = NULL,
  adjusted = NULL,
  multi = NULL,
  effect_ref = NULL,
  template = c("journal", "clean", "minimal"),
  append = NULL,
  file = NULL,
  raw = FALSE,
  name = FALSE,
  title = NULL,
  show = TRUE,
  mode = c("auto", "console", "table")
)

Arguments

...

In console mode, one row variable and optionally one column variable, followed by console options such as exp, chi, and fisher. In publication mode, legacy positional data, vars, and by arguments are also accepted.

data

Optional data frame. When omitted or NULL, the active data frame set by usedf() or opendata(..., active = TRUE) is used.

vars

A variable specification created by vars().

by

Optional grouping or outcome variable supplied without quotation marks. Leave it empty for an overall descriptive table. Use a regular variable name for a categorical grouping/outcome variable, c.outcome for a continuous outcome summarized by mean (SD), or q.outcome for a continuous outcome summarized by median (IQR). R4VN also accepts the unified hierarchical form by = vars(province, sex, outcome): province and sex are nested superby strata, in that order, and outcome is the innermost grouping/outcome variable.

superby

Backward-compatible single stratification variable. It may be combined with hierarchical by = vars(...) and then becomes the outermost stratum. New code should normally prefer the unified by convention.

digit

Number of decimal places for descriptive statistics.

p_digit

Number of decimal places for p-values.

effect_digit

Number of decimal places for OR, RR, PR, or linear regression coefficients.

missing

Missing-value display for categorical variables: "no", "ifany", or "always".

row

Logical. Calculate row percentages when by is categorical. When row = TRUE, col and cell are automatically set to FALSE.

col

Logical. Calculate column percentages when by is categorical. This is the default percentage mode.

cell

Logical. Calculate percentages using the complete table total. When cell = TRUE and row = FALSE, row and col are automatically set to FALSE. Thus users normally need to specify only row = TRUE, cell = TRUE, or neither for the default column percentages.

overall

Position of the overall column: "none", "first", or "last". Logical values are accepted for backward compatibility.

descriptive

Logical. Display descriptive-statistics columns.

rvrow

Categorical variables whose displayed level order should be reversed. Accepts TRUE, vars(...), c(...), a single variable name, or a character vector. This does not change model reference categories.

rvcol

Logical. Reverse displayed levels of a categorical by variable.

test

Logical. Display traditional omnibus-test p-values.

pvalue

Logical. Display separate p-value columns for model coefficients.

bold_p

Logical. Bold p-values smaller than p_bold.

p_bold

Significance threshold used when bold_p = TRUE.

test_note

Logical. Add superscript letters and footnotes identifying omnibus tests.

interaction

Logical. When superby is supplied, add one final interaction p-value column. Interaction tests use predictor-by-superby terms and follow multi, then adjusted, then crude models.

or

Logical. Calculate odds ratios using logistic regression.

rr

Logical. Calculate risk ratios using modified Poisson regression with robust variance.

pr

Logical. Calculate prevalence ratios using modified Poisson regression with robust variance.

event

Event level of a binary outcome. The last observed level is used when omitted.

adjusted

Variables included as adjustment covariates in separate models for each focal predictor. Prefer vars(c.age, b2.sex, q.bmi) so variable types and reference levels remain explicit. Also accepts c(...), a character vector, TRUE, or "ALL".

multi

Variables included together in one final multivariable model. Prefer vars(c.age, b2.sex, c.bmi). TRUE or "ALL" includes every variable listed in vars.

effect_ref

Optional backward-compatible reference categories. The b2., b3., and related prefixes take precedence.

template

HTML style: "journal", "clean", or "minimal".

append

Optional previous r4vn_tab object or existing HTML path.

file

Optional output HTML path. A temporary file is created when omitted.

raw

Logical. Retain unformatted results in the returned object.

name

Logical. Display original variable names beside variable labels.

title

Optional table title.

show

Logical. Display the HTML table in the RStudio Viewer or browser.

mode

Dispatch mode. "auto" selects publication mode when a vars() specification is supplied and otherwise selects console mode. Use "console" or "table" to force a mode.

Details

Prefixes used inside vars() determine descriptive summaries and categorical reference levels:

  • no prefix: automatic typing; numeric/integer variables use mean and standard deviation, while factor/character/logical variables are categorical with the first observed level as reference;

  • b1., b2., b3., ...: force a categorical variable with the corresponding observed level as reference;

  • c.: mean and standard deviation;

  • q.: median and interquartile range;

  • f.: mean, median, and range.

For grouped categorical tables, column percentages are the default. Setting row = TRUE automatically turns col and cell off; setting cell = TRUE automatically turns row and col off. Users therefore do not need to manually disable col = TRUE.

With a categorical by variable, categorical predictors are tested using Pearson's chi-squared test or Fisher's exact test. Variables declared with c. use a t-test or one-way ANOVA; variables declared with q. or f. use the Wilcoxon rank-sum or Kruskal-Wallis test.

Binary outcomes can be analyzed with OR, RR, or PR. OR uses logistic regression. RR and PR use modified Poisson regression with robust variance.

With by = c.outcome, the continuous outcome is summarized by mean (SD), categorical predictors use t-tests/ANOVA, and numeric predictors use Pearson correlation tests. With by = q.outcome, the outcome is summarized by median (IQR), categorical predictors use Wilcoxon/Kruskal-Wallis tests, and numeric predictors use Spearman tests. Both modes report unstandardized beta coefficients from linear regression.

adjusted and multi have different roles. adjusted fits a separate adjusted model for each focal predictor. multi fits one final model containing all specified variables.

When superby is supplied, tab() first calculates the complete dataset and then repeats the same analysis independently within every level of superby. The resulting blocks are combined side by side. When possible, one final interaction p-value column tests whether each predictor effect differs across the levels of superby.

Value

Invisibly returns an object of class r4vn_tab. Important components include data, file, html, table_html, rows, multi_model, and multi_diagnostics.

Common call patterns

tab(data, vars = vars(...))
tab(data, vars = vars(...), by = group)
tab(data, vars = vars(...), by = outcome, or = TRUE)
tab(data, vars = vars(...), by = c.outcome)
tab(data, vars = vars(...), by = q.outcome)
tab(data, vars = vars(...), by = outcome, superby = subgroup, or = TRUE)

See Also

vars, tabmulti, and tabexport.

Other R4VN tables: tabexport(), tabforest(), tablong(), tabmeta(), tabmulti(), tabscale(), tabscore(), tabsurvey(), vars()

Examples

set.seed(2026)
n <- 180
dat <- data.frame(
  age = round(rnorm(n, 45, 12)),
  sex = factor(sample(c("Female", "Male"), n, TRUE)),
  bmi = round(rnorm(n, 23, 3), 1),
  smoking = factor(sample(c("No", "Yes"), n, TRUE,
                          prob = c(0.70, 0.30))),
  education = factor(sample(c("Primary", "Secondary", "College"),
                            n, TRUE))
)
dat$sbp <- round(80 + 0.75 * dat$age + 1.1 * dat$bmi +
                 5 * (dat$sex == "Male") +
                 4 * (dat$smoking == "Yes") + rnorm(n, 0, 12), 1)
lp <- -3.2 + 0.045 * dat$age + 0.10 * (dat$bmi - 23) +
      0.45 * (dat$sex == "Male") + 0.65 * (dat$smoking == "Yes")
dat$hypertension <- factor(
  rbinom(n, 1, plogis(lp)),
  levels = c(0, 1), labels = c("No", "Yes")
)

tb0 <- tab(dat, vars = vars(c.age, b2.sex, q.bmi, b2.smoking, education),
           show = FALSE)
head(tb0$data)

tb1 <- tab(dat, vars = vars(c.age, b2.sex, c.bmi, b2.smoking, b2.education),
           by = hypertension, or = TRUE, event = "Yes",
           multi = vars(c.age, b2.sex, c.bmi, b2.smoking), show = FALSE)

tb2 <- tab(dat, vars = vars(c.age, b2.sex, c.bmi, b2.smoking),
           by = c.sbp, multi = vars(c.age, b2.sex, c.bmi, b2.smoking),
           show = FALSE)

tb3 <- tab(dat, vars = vars(c.age, c.bmi, b2.smoking, b2.education),
           by = hypertension, superby = sex, overall = "none",
           or = TRUE, event = "Yes",
           multi = vars(c.age, c.bmi, b2.smoking), show = FALSE)

# Extended usage examples

set.seed(2026)
n <- 300
d <- data.frame(
  sex = factor(sample(c("Female", "Male"), n, TRUE)),
  age = rnorm(n, 45, 12),
  bmi = rnorm(n, 23, 3),
  smoking = factor(sample(c("No", "Yes"), n, TRUE)),
  region = factor(sample(c("Urban", "Rural"), n, TRUE)),
  outcome = factor(rbinom(n, 1, .3), levels = 0:1, labels = c("No", "Yes")),
  sbp = rnorm(n, 125, 18)
)

# Overall descriptive table. Numeric variables without a prefix are
# automatically summarized with mean (SD); factors remain categorical.
t1_auto <- tab(d, vars = vars(age, sex, bmi, smoking), show = FALSE)

# Explicit q. remains available when median (IQR) is preferred.
t1 <- tab(d, vars = vars(sex, age, q.bmi, smoking), show = FALSE)

# Compare groups, show overall first, tests, and missing values when present
t2 <- tab(d, vars = vars(sex, c.age, q.bmi, smoking), by = outcome,
          overall = "first", test = TRUE, missing = "ifany", show = FALSE)

# Row, column, or cell percentages for categorical variables
tab(d, vars = vars(sex, smoking), by = outcome, row = TRUE, show = FALSE)
tab(d, vars = vars(sex, smoking), by = outcome, show = FALSE)
tab(d, vars = vars(sex, smoking), by = outcome, cell = TRUE, show = FALSE)

# Reverse selected row levels or the by-variable columns
tab(d, vars = vars(sex, smoking), by = outcome,
    rvrow = vars(smoking), rvcol = TRUE, show = FALSE)

# Crude odds ratios for a binary outcome
tab(d, vars = vars(sex, c.age, smoking), by = outcome,
    or = TRUE, event = "Yes", show = FALSE)

# Risk ratios or prevalence ratios using modified Poisson models
tab(d, vars = vars(sex, c.age, smoking), by = outcome,
    rr = TRUE, event = "Yes", show = FALSE)
tab(d, vars = vars(sex, c.age, smoking), by = outcome,
    pr = TRUE, event = "Yes", show = FALSE)

# Separate adjusted models for every focal predictor
tab(d, vars = vars(sex, c.age, smoking), by = outcome,
    or = TRUE, adjusted = vars(age, sex), event = "Yes", show = FALSE)

# One final multivariable model; effects are placed beside their variables
tab(d, vars = vars(sex, c.age, smoking, q.bmi), by = outcome,
    or = TRUE, multi = vars(sex, age, smoking), event = "Yes", show = FALSE)

# Hide descriptive columns and show only model results
tab(d, vars = vars(sex, c.age, smoking), by = outcome,
    descriptive = FALSE, or = TRUE, multi = TRUE,
    event = "Yes", show = FALSE)

# Continuous outcome: c. gives parametric methods and beta coefficients
tab(d, vars = vars(sex, c.age, smoking, q.bmi), by = c.sbp,
    adjusted = vars(age, sex), multi = vars(age, sex, bmi), show = FALSE)

# Continuous outcome: q. gives rank-based descriptive comparisons
tab(d, vars = vars(sex, c.age, smoking, q.bmi), by = q.sbp,
    test = TRUE, show = FALSE)

# Supergroup columns plus interaction
tab(d, vars = vars(sex, c.age, smoking), by = outcome, superby = region,
    interaction = TRUE, overall = "first", show = FALSE)

# Templates, titles, raw numerical output, and named variables
tab(d, vars = vars(sex, c.age, smoking), by = outcome,
    template = "minimal", title = "Participant characteristics",
    raw = TRUE, name = TRUE, show = FALSE)


R4VN documentation built on Sept. 30, 2026, 5:13 p.m.