tabforest: Flexible regression, multi-outcome, survival, and subgroup...

View source: R/tabforest.R

tabforestR Documentation

Flexible regression, multi-outcome, survival, and subgroup forest plots

Description

tabforest() is the common forest-plot engine for R4VN. It can (1) fit regression models directly from an outcome and focal predictors, (2) reuse a fitted R4VN or standard R model, (3) place several outcomes side by side using the same predictor structure, and (4) create subgroup-effect forests with a p-value for interaction. Numeric estimates are always stored without clipping; xmin and xmax affect only the drawing.

Usage

tabforest(
  outcome = NULL,
  predictors = NULL,
  data = NULL,
  time = NULL,
  event = NULL,
  failure = NULL,
  outcomes = NULL,
  subgroup = NULL,
  predictor = NULL,
  type = c("auto", "regression", "multioutcome", "subgroup"),
  crude = TRUE,
  adjusted = FALSE,
  multi = FALSE,
  or = FALSE,
  rr = FALSE,
  pr = FALSE,
  irr = FALSE,
  estimate = c("auto", "beta", "or", "rr", "pr", "irr", "hr"),
  ci = 0.95,
  sample = c("auto", "common", "model"),
  per = NULL,
  per_labels = NULL,
  select = NULL,
  xmin = NULL,
  xmax = NULL,
  ticks = NULL,
  log = NULL,
  arrows = TRUE,
  row_layout = c("auto", "modelrows", "compact"),
  layout = c("dodge", "stack"),
  row_spacing = 1,
  model_row_gap = 0.55,
  group_gap = 0.25,
  order = NULL,
  reference = TRUE,
  pvalue = TRUE,
  global_p = FALSE,
  show_n = FALSE,
  show_events = FALSE,
  show_model_label = TRUE,
  show_interaction_p = TRUE,
  p_layout = c("inline", "column"),
  effect_digit = 2,
  p_digit = 3,
  labels = NULL,
  level_labels = NULL,
  lang = c("en", "vi"),
  text = NULL,
  model_labels = NULL,
  label_title = NULL,
  effect_title = NULL,
  axis_title = NULL,
  title = NULL,
  subtitle = NULL,
  caption = NULL,
  note = TRUE,
  template = c("journal", "clean", "minimal"),
  grid = c("major", "none", "both"),
  font_family = "",
  base_size = 11,
  label_cex = 1,
  header_cex = 1,
  axis_cex = 1,
  model_cex = 0.86,
  colors = NULL,
  fills = NULL,
  pch = NULL,
  lty = NULL,
  point_cex = 1.15,
  point_lwd = 1,
  ci_lwd = 1.2,
  ref_lwd = 1,
  ref_lty = 2,
  ref_col = "gray45",
  arrow_length = 0.08,
  zebra = FALSE,
  zebra_fill = c("white", "gray94"),
  zebra_by = c("variable", "header"),
  label_width = 0.28,
  forest_width = 0.32,
  column_gap = 0.012,
  panel_gap = 0.012,
  panel_forest_ratio = 0.58,
  label_indent = 0.018,
  label_wrap = 38,
  file = NULL,
  width = 12,
  height = NULL,
  dpi = 300,
  show = TRUE,
  console = FALSE
)

Arguments

outcome

Outcome variable for ordinary regression/subgroup analysis, or a supported fitted object (r4vn_surv, r4vn_tabmulti, r4vn_stat, lm, glm, coxph). For Cox regression, outcome is the event/status variable. May be omitted when outcomes is supplied.

predictors

Focal predictors that should appear in the forest. Prefer vars() so R4VN type/reference declarations are retained, for example vars(c.age, b2.sex, c.bmi, smoking).

data

Data frame. If omitted, the active R4VN data frame is used.

time

Follow-up time variable for Cox regression. Supplying time automatically selects HR unless another incompatible effect is requested.

event

Modeled event level for binary OR/RR/PR outcomes.

failure

Event value for Cox regression. With numeric 0/1 status, 1 is selected automatically when present.

outcomes

Optional named character vector or named list for a multi-outcome forest. Character example: c(HTN="hypertension", DEP="depression"). A list allows outcome-specific settings, e.g. list(HTN=list(outcome="hypertension",event="Yes"), Death=list(outcome="death",time="followup",failure=1,estimate="hr")).

subgroup

Optional categorical variables for subgroup analysis. When supplied, use predictor for the main exposure whose effect is estimated within each subgroup level.

predictor

Main exposure for subgroup mode. It may be continuous or a two-level categorical variable. Use vars(b2.treatment) when a non-default reference is required.

type

Analysis mode: "auto", "regression", "multioutcome", or "subgroup". "auto" chooses multi-outcome when outcomes is non-NULL, subgroup when subgroup is non-NULL, otherwise ordinary regression.

crude

For regression/multi-outcome mode, fit one crude model per focal predictor. Set FALSE when only adjusted/multivariable estimates are wanted.

adjusted

Regression mode: FALSE, TRUE, or vars(...). With vars(X1,X2), each focal predictor gets a separate model adjusted for X1/X2. With TRUE, each focal predictor is adjusted for all other focal predictors. Subgroup mode requires an explicit vars(...) adjustment set if adjustment is desired.

multi

Regression mode: FALSE, TRUE, or vars(...). TRUE fits one joint model containing all focal predictors. vars(A,B,C,X) fits that exact joint model but still displays only variables listed in predictors. In subgroup mode, an explicit vars(...) may also be used as the final adjustment set; the exposure and current subgroup variable are removed from the covariate set automatically.

or, rr, pr, irr

Logical shortcuts for OR, RR, PR, or IRR. Only one may be TRUE. Binary outcomes default to OR; numeric outcomes default to beta.

estimate

Explicit effect type: "auto", "beta", "or", "rr", "pr", "irr", or "hr".

ci

Confidence level, default 0.95.

sample

Missing-data strategy. "common" forces displayed regression models to use the same complete-case sample; "model" allows each model to use its own available cases; "auto" uses a common sample when several regression model groups are displayed. In subgroup mode, "auto" behaves like model-specific analysis so unrelated subgroup variables do not reduce one another's sample size.

per

Optional multiplier for continuous effects. Example c(age=10) reports the ratio/HR per 10 years or beta per 10 units.

per_labels

Optional display labels for per, e.g. c(age="per 10 years").

select

Model components when outcome is a fitted R4VN object. For tabsurv() this can include "crude", "adjusted", "multi"; for tabmulti() use stored model names such as "full", "backward".

xmin, xmax

Forest plotting limits. These never alter stored estimates. For ratio effects, if only xmax is supplied and is >1, xmin=1/xmax is used automatically. In multi-outcome mode these may be named vectors or lists keyed by outcome-panel name.

ticks

Optional axis ticks. In multi-outcome mode a named list can give different ticks to different outcome panels.

log

NULL or logical. Ratio effects default to logarithmic axes, beta to linear. In multi-outcome mode this may be a named logical vector/list.

arrows

Draw arrowheads when CIs extend beyond plotting limits. When the point estimate itself is outside the range, no false boundary point is drawn.

row_layout

"auto", "modelrows", or "compact". The default "auto" uses "modelrows" whenever more than one model group is displayed and "compact" when only one model is displayed. In "modelrows", Crude, Adjusted, and Multivariable are separate physical rows but share ONE effect column (for example, one ⁠OR (95% CI)⁠ column). Thus each CI, marker, numeric estimate, and p-value is aligned with its own row. Use "compact" only when several model estimates are deliberately wanted on the same labelled row.

layout

In compact mode, "dodge" or "stack" controls vertical offsets of multiple model markers/CI lines.

row_spacing

Baseline distance between ordinary rows.

model_row_gap

Distance between model rows belonging to the same variable/level when row_layout="modelrows".

group_gap

Extra vertical separation between variable blocks.

order

Optional order of focal predictor variable names.

reference

Show categorical reference rows. Default TRUE.

pvalue

Show coefficient-level p-values.

global_p

Show categorical-variable omnibus Wald p-values. Default FALSE because forest plots are usually cleaner without these values.

show_n, show_events

Add model N and number of events after the estimate.

show_model_label

In modelrows, print the model label (Crude, Adjusted, Multivariable) beside the corresponding row.

show_interaction_p

In subgroup mode, show the p-value for interaction on the subgroup-variable header row.

p_layout

Regression display style: "inline" appends p to the numeric effect string; "column" uses a separate p-value column. modelrows works particularly well with either style.

effect_digit, p_digit

Decimal places for effects and p-values.

labels

Named character vector overriding variable labels.

level_labels

Named list overriding displayed categorical levels. This changes display only, not model coding/reference levels.

lang

Built-in language: "en" or "vi".

text

Named list overriding individual words. Useful keys include characteristic, reference, crude, adjusted, multi, p, n, events, subgroup, interaction_p, overall, arrow_note, and effect keys OR, RR, PR, IRR, HR, beta.

model_labels

Named character vector overriding model labels, e.g. c(Crude="Unadjusted",Multivariable="Adjusted").

label_title, effect_title, axis_title

Optional column/axis titles.

title, subtitle, caption

Optional plot title, subtitle, caption.

note

TRUE for an automatic clipping note, FALSE for none, or custom text.

template

Visual preset: "journal", "clean", "minimal".

grid

"major", "none", or "both".

font_family

Base graphics font family.

base_size, label_cex, header_cex, axis_cex, model_cex

Text-size controls.

colors

Model line/marker colors. A named vector is recommended, e.g. c(Crude="gray50",Adjusted="navy",Multivariable="firebrick").

fills

Optional marker fill colors, useful with pch 21:25.

pch

Model marker symbols. Named vectors may assign different symbols to crude and adjusted estimates.

lty

Model CI line types.

point_cex

Model marker sizes. May be scalar or named vector by model.

point_lwd

Marker border widths. May be scalar or named vector.

ci_lwd

CI line widths. May be scalar or named vector by model.

ref_lwd, ref_lty, ref_col

Null-line appearance.

arrow_length

Arrowhead size in inches.

zebra

Draw alternating background blocks by predictor/subgroup.

zebra_fill

Two or more background colors, e.g. c("white","gray93").

zebra_by

"variable" or "header"; both currently alternate complete variable/subgroup blocks so all levels/models in a block share a background.

label_width, forest_width, column_gap

Horizontal layout controls for a single regression/subgroup forest.

panel_gap

Gap between panels in multi-outcome mode.

panel_forest_ratio

Fraction of each multi-outcome panel devoted to the CI forest; the remaining panel width is used for numeric estimates.

label_indent

Indentation of categorical levels.

label_wrap

Approximate wrapping width for long labels; Inf disables.

file

Optional PDF/PNG/SVG/JPG/TIFF output file.

width, height, dpi

Graphics dimensions. If height is NULL it grows with the actual number of drawn rows, so modelrows can produce a tall figure without compressing row spacing.

show

Draw immediately. Default TRUE.

console

Print the long standardized estimate table.

Details

1. Regression model semantics

With predictors=vars(A,B,C):

  • crude=TRUE: Y~A, Y~B, Y~C.

  • adjusted=vars(X): Y~A+X, Y~B+X, Y~C+X.

  • adjusted=TRUE: each focal variable is adjusted for the other focal vars.

  • multi=TRUE: one joint model Y~A+B+C.

  • multi=vars(A,B,C,X): one joint model Y~A+B+C+X, but only A/B/C are shown.

Therefore adjusted and multi answer different scientific questions and may be requested together in the same forest. When two or more model groups are displayed, row_layout="auto" uses separate model rows and ONE shared effect column. For example, Crude and Multivariable ORs are both printed under the same ⁠OR (95% CI)⁠ header instead of being put in separate Crude-OR and Multivariable-OR columns.

2. Effect measure selected by outcome

  • numeric continuous outcome -> linear regression -> beta, null=0;

  • binary outcome -> logistic regression -> OR, null=1;

  • pr=TRUE -> robust modified Poisson -> PR, null=1;

  • rr=TRUE -> robust modified Poisson -> RR, null=1;

  • count outcome + irr=TRUE -> Poisson -> IRR, null=1;

  • ⁠time=⁠ -> Cox proportional hazards -> HR, null=1.

3. Exact numbers when the forest is clipped

Suppose OR=7.41 and 95% CI=1.56 to 35.20 while xmax=10. The printed number remains 7.41 (1.56-35.20). Only the graphical CI is truncated at 10 and an arrow is drawn. If OR itself exceeds 10, no marker is placed falsely at 10.

4. Multiple models: separate rows, one merged effect column

row_layout="auto" is the default. If more than one model group is present, it automatically switches to the "modelrows" layout. Crude, Adjusted and/or Multivariable estimates are placed on separate physical rows, while the right side contains only ONE shared effect column such as ⁠OR (95% CI)⁠, ⁠HR (95% CI)⁠, or ⁠Beta (95% CI)⁠. The result, CI line, marker and p-value therefore stay on exactly the same row. This is the recommended publication layout when crude and adjusted estimates are presented together. Increase row_spacing, model_row_gap, group_gap, or leave height=NULL for a taller figure.

Set row_layout="compact" only when you intentionally want several model estimates on the same labelled row; compact mode retains separate numeric model columns because the rows are not expanded.

5. Multi-outcome forests

⁠outcomes=⁠ creates side-by-side panels sharing predictor labels. Every panel may have its own effect type, axis range and follow-up variable. This permits two binary outcomes (OR panels), several continuous outcomes (beta panels), or even mixed OR/HR panels in one figure. For very many panels, increase width.

6. Subgroup forests

⁠subgroup=⁠ estimates the effect of one main predictor separately within each subgroup level. A full model containing predictor*subgroup is fitted for the Wald interaction p-value. For RR/PR, the interaction Wald test uses the same robust covariance approach as the effect model. Cox subgroup forests use HR and a Cox interaction model. A subgroup variable must be categorical; make clinically meaningful groups before calling tabforest().

7. Styling

colors, fills, pch, lty, point_cex, point_lwd, and ci_lwd accept named model vectors. zebra=TRUE shades complete variable blocks, closely matching journal forest-table layouts. lang="vi", ⁠text=⁠, ⁠labels=⁠, and ⁠level_labels=⁠ allow all visible wording to be translated without changing the model.

Value

An object of class r4vn_tabforest; multi-outcome and subgroup modes add subclasses r4vn_tabforest_multi and r4vn_tabforest_subgroup. ⁠$data⁠ is publication-ready, ⁠$table⁠ is the numeric long table, ⁠$models⁠ stores fitted models, and plot() can redraw without refitting.

See Also

vars, tab, tabmulti, tabsurv, tabexport

Other R4VN tables: tab(), tabexport(), tablong(), tabmeta(), tabmulti(), tabscale(), tabscore(), tabsurvey(), vars()

Examples

data(tabforest_demo)
usedf(tabforest_demo, quiet = TRUE)

# Crude odds ratios.
f1 <- tabforest(
  hypertension,
  predictors = vars(c.age, sex, c.bmi, smoking),
  event = "Yes",
  show = FALSE
)
f1$data

# Crude plus one final multivariable model.
f2 <- tabforest(
  hypertension,
  predictors = vars(c.age, sex, c.bmi, smoking),
  event = "Yes",
  crude = TRUE, multi = TRUE,
  show = FALSE
)

# Re-drawing is intentionally interactive so CRAN examples do not depend
# on the graphics device or installed fonts.
if (interactive()) {
  plot(f2, row_layout = "modelrows", zebra = TRUE)
}

# Each focal predictor adjusted for the same confounders.
f3 <- tabforest(
  hypertension,
  predictors = vars(c.age, c.bmi, smoking),
  event = "Yes",
  adjusted = vars(sex, education),
  show = FALSE
)

# Vietnamese display text can be prepared without drawing during checks.
f_vi <- tabforest(
  hypertension,
  predictors = vars(c.age, sex, c.bmi, smoking),
  event = "Yes",
  multi = TRUE,
  lang = "vi",
  labels = c(
    age = "Tu\u1ed5i",
    sex = "Gi\u1edbi t\u00ednh",
    bmi = "Ch\u1ec9 s\u1ed1 kh\u1ed1i c\u01a1 th\u1ec3",
    smoking = "H\u00fat thu\u1ed1c"
  ),
  level_labels = list(
    sex = c(Female = "N\u1eef", Male = "Nam"),
    smoking = c(No = "Kh\u00f4ng", Yes = "C\u00f3")
  ),
  text = list(reference = "Tham chi\u1ebfu"),
  title = "Bi\u1ec3u \u0111\u1ed3 forest",
  show = FALSE
)
if (interactive()) plot(f_vi)


# Modified-Poisson prevalence ratio.
f_pr <- tabforest(
  depression,
  predictors = vars(c.age, sex, smoking, alcohol),
  event = "Yes", pr = TRUE, multi = TRUE,
  show = FALSE
)

# Continuous outcome.
f_beta <- tabforest(
  sbp,
  predictors = vars(c.age, sex, c.bmi, smoking),
  multi = TRUE,
  show = FALSE
)

# Cox model, only when the suggested package is available.
if (requireNamespace("survival", quietly = TRUE)) {
  f_hr <- tabforest(
    death, time = followup,
    predictors = vars(c.age, sex, treatment, c.bmi),
    failure = 1, crude = TRUE, multi = TRUE,
    show = FALSE
  )
}

# Multi-outcome forest without drawing.
f_multi <- tabforest(
  outcomes = c(
    Hypertension = "hypertension",
    Depression = "depression"
  ),
  predictors = vars(c.age, sex, c.bmi, smoking),
  event = "Yes", crude = FALSE, multi = TRUE,
  show = FALSE
)

# Subgroup forest without drawing.
f_sub <- tabforest(
  hypertension,
  predictor = vars(treatment),
  subgroup = vars(age_group, sex, obesity, diabetes, smoking),
  event = "Yes", type = "subgroup",
  adjusted = vars(c.age, c.bmi),
  show = FALSE
)

# File output uses a temporary path and is cleaned up.
f_png <- tempfile(fileext = ".png")
tabforest(
  hypertension,
  predictors = vars(c.age, sex, c.bmi, smoking),
  event = "Yes", multi = TRUE,
  file = f_png, width = 8, height = 5, dpi = 120,
  show = FALSE
)
unlink(f_png)


usedf(clear = TRUE, quiet = TRUE)

R4VN documentation built on Sept. 30, 2026, 5:13 p.m.