tablong: Longitudinal and Repeated-Measures Analysis

tablongR Documentation

Longitudinal and Repeated-Measures Analysis

Description

Performs publication-ready longitudinal or repeated-measures analysis from either long or wide data. tablong() is designed for the usual biomedical workflow: describe each time point, test overall time and group effects, test the time-by-group interaction, estimate clinically interpretable contrasts, retain fitted models for advanced use, and optionally create a longitudinal profile plot and a cautious interpretation table.

Usage

tablong(
  data = NULL,
  vars,
  time = NULL,
  id = NULL,
  by = NULL,
  ref = NULL,
  event = NULL,
  adjusted = NULL,
  gee = FALSE,
  ar1 = FALSE,
  slope = FALSE,
  or = FALSE,
  rr = FALSE,
  pr = FALSE,
  count = FALSE,
  exposure = NULL,
  change = TRUE,
  pairwise = FALSE,
  adjust = "none",
  missing = FALSE,
  level = 0.95,
  digit = 1,
  p_digit = 3,
  effect_digit = 2,
  bold_p = TRUE,
  p_bold = 0.05,
  diagnostics = FALSE,
  diagnosis = FALSE,
  interpretation = FALSE,
  plot = FALSE,
  plot_args = list(),
  name = FALSE,
  title = NULL,
  file = NULL,
  raw = FALSE,
  show = TRUE
)

Arguments

data

Optional data frame. If NULL, the active data selected by usedf() are used.

vars

Outcome specification created by vars(). In long data, several outcomes can be analyzed in one call. Continuous outcomes use the c. prefix, q. requests median (IQR) descriptive display, and f. requests a fuller continuous summary. An unprefixed two-level variable is treated as binary. In wide data, the variables in vars() are repeated measurements of the same outcome.

time

In long data, an unquoted time variable. Use c.timevar for a continuous linear time effect or b2.timevar, b3.timevar, etc. to select the categorical reference level. In wide data, provide display labels such as c("Baseline", "Month 3", "Month 6"); when omitted, the repeated variable names are used as time labels.

id

Optional subject identifier. In wide data it is optional because each source row represents one subject. In long data, repeated IDs trigger longitudinal analysis with within-subject correlation. If no ID is supplied, or IDs do not repeat across time, observations are analyzed as repeated cross-sectional samples.

by

Optional grouping variable, for example treatment group. Prefix b2., b3., etc. selects the reference group.

ref

Optional categorical reference time label. This overrides a bN. prefix supplied in time.

event

Event level for binary outcomes. A single value applies to every binary outcome; a named vector can specify a different event for each outcome, for example event = c(controlled = "Yes", admitted = "Yes").

adjusted

Optional adjustment variables created by vars(). Use the same R4VN prefixes as elsewhere, for example vars(c.age, sex, b2.site).

gee

Logical. Request a population-average marginal model for repeated data. By default R4VN uses working-independence regression with subject- clustered robust sandwich standard errors calculated internally. No extra package is required. Set ar1 = TRUE to request AR(1) GEE through optional package geepack when it is installed.

ar1

Logical. Request an AR(1) working correlation for a marginal GEE. This advanced option uses optional package geepack. If it is unavailable, tablong() warns and falls back to working-independence cluster-robust inference instead of stopping the analysis.

slope

Logical. With a mixed model and continuous time (time = c.month), include a subject-specific random linear time slope in addition to the random intercept.

or, rr, pr

Logical effect switches for binary outcomes. Odds ratio is the default when none is selected. rr = TRUE reports risk ratios and pr = TRUE reports prevalence ratios using modified Poisson regression with robust variance; repeated subjects use subject-clustered robust variance. No external sandwich/GEE package is needed unless ar1 = TRUE is requested. Only one switch may be TRUE.

count

Logical. Treat numeric outcomes as non-negative counts and fit Poisson models. Count outcomes report incidence-rate ratios (IRR).

exposure

Optional positive exposure/person-time variable for count models. In wide data it may also be vars(exp0, exp1, ...), with one exposure variable per repeated count variable.

change

Logical. For categorical time, report change from the reference time. With exactly two groups, also report the difference in change, i.e. the usual difference-in-differences contrast. Default TRUE.

pairwise

Logical. Calculate all available time and group pairwise contrasts and retain them in ⁠$contrasts⁠ and ⁠$contrasts_table⁠. The main publication table remains compact. Default FALSE.

adjust

Multiplicity adjustment applied to contrast p-values. Any method accepted by p.adjust() may be used, including "none", "holm", "bonferroni", and "BH".

missing

Logical. Append cell-specific n to continuous/count summary cells. Binary cells always show event/total. Detailed observed and missing counts are always available in ⁠$descriptive⁠.

level

Confidence level, default 0.95.

digit

Decimal places for descriptive summaries.

p_digit

Decimal places for p-values.

effect_digit

Decimal places for model effects and confidence intervals.

bold_p

Logical. Bold p-values below p_bold in the HTML Viewer.

p_bold

Threshold used by bold_p.

diagnostics

Logical. Include the compact model-diagnostics table in the HTML Viewer. Diagnostics are always retained in ⁠$diagnostics⁠; default FALSE keeps the primary Viewer concise.

diagnosis

Logical singular alias for diagnostics, provided for consistency with other R4VN regression commands. Default FALSE. When explicitly supplied, it overrides diagnostics; omitting it preserves backward-compatible use of diagnostics.

interpretation

Logical. Add a cautious deterministic interpretation table. The default is FALSE. The interpretation emphasizes the time-by-group interaction when present and does not replace substantive or clinical interpretation by the researcher.

plot

Logical. Create an observed longitudinal profile plot with 95% confidence intervals, include the same plot directly in the HTML Viewer, display it in the Plot pane, and store its specification in ⁠$graph⁠ and ⁠$plots$trajectory⁠. Plotting uses base R graphics; ggplot2 is not required. Default FALSE.

plot_args

Named list controlling the profile plot. Supported entries include title, xlab, ylab, ci, line_width, point_size, base_size, legend_position, font_family, colors, point_shapes, line_types, and grid. Generic sans is the default font for reliable display in RStudio Viewer, browsers, Windows, macOS, and Linux.

name

Logical. Display the original variable name after its variable label in the main table.

title

Optional table title.

file

Optional HTML file path. When omitted, a temporary HTML file is created. This file is the formatted Viewer report, not a replacement for tabexport().

raw

Logical retained for backward compatibility. Raw models, tests, contrasts, standardized long data, and reporting tables are always retained in the returned object.

show

Logical. Open the formatted HTML report in the Viewer/browser. Default TRUE.

Details

The interface follows the R4VN principle of keeping routine analysis simple. In most studies the essential call is only vars(), time, id, and optionally by. Repeated continuous outcomes use a random-intercept model when R's recommended nlme package is available. Repeated binary/count outcomes use marginal regression with subject-clustered robust standard errors calculated internally by R4VN, so routine analyses need no extra package.

Data format. If time names a column in data, input is treated as long. If vars() contains multiple repeated variables and time is a vector of labels (or omitted), input is treated as wide and is reshaped internally. The original data frame is never modified.

Continuous outcomes. Repeated subjects use a random-intercept linear mixed model through R's recommended nlme package when available. slope = TRUE adds a random linear time slope when time is continuous. If nlme is unavailable, R4VN falls back to a marginal linear model with subject-clustered robust standard errors. Repeated cross-sectional data use ordinary linear models. gee = TRUE explicitly requests the marginal model.

Binary outcomes. The default effect is an odds ratio from logistic regression. For repeated subjects, R4VN calculates subject-clustered robust standard errors internally. rr = TRUE and pr = TRUE use modified Poisson regression with robust variance and report RR or PR. This avoids requiring lme4, sandwich, or geepack for routine binary longitudinal analysis.

Count outcomes. count = TRUE fits a Poisson model and reports IRR. Repeated subjects use subject-clustered robust standard errors calculated internally. Supplying exposure adds offset(log(exposure)) and the observed plot is an incidence rate per one person-time unit.

Categorical time. The model includes time, group when supplied, and the time-by-group interaction. The main table reports observed summaries at each time. With two groups it also reports the between-group effect at each time, within-group change from the reference time, and the difference in change. The omnibus ⁠Time x group⁠ p-value is the formal test that temporal changes differ between groups.

Continuous time. time = c.month estimates change per one time unit. With a group variable, group-specific slopes and their difference are returned.

More than two groups. The main table remains intentionally compact and shows omnibus tests. Set pairwise = TRUE to obtain all model-based pairwise comparisons in ⁠$contrasts_table⁠; use adjust to control multiplicity.

Descriptive prefixes. q. and f. affect the observed descriptive summary only. Inferential effects remain based on the selected mean model; they do not fit median regression.

Missing values. Each model uses observations complete for that outcome, time, ID/group, exposure if required, and adjustment variables. Descriptive counts are retained separately so that missingness and attrition can be reviewed before publication.

Package dependencies. Routine tablong() analyses and plots are designed to run with base/recommended R only. nlme is used for continuous random- effects models and is bundled with standard R installations. geepack is optional and used only when an AR(1) GEE is explicitly requested. lme4, sandwich, broom, and ggplot2 are not required by tablong().

Returned reporting contract. For programmatic reuse, the object includes a flat main table plus descriptive, omnibus-test, contrast, diagnostics, interpretation, plot, model, and metadata components. The object also inherits from r4vn_tab, so existing tabexport() workflows continue to work.

Value

Invisibly returns an object of class c("r4vn_tablong", "r4vn_tab"). Important components are: ⁠$data⁠ (main publication table), ⁠$descriptive⁠, ⁠$tests⁠ and ⁠$tests_table⁠, ⁠$contrasts⁠ and ⁠$contrasts_table⁠, ⁠$diagnostics⁠, ⁠$interpretation⁠, ⁠$tables⁠, ⁠$graph⁠, ⁠$plots⁠, ⁠$models⁠, ⁠$long_data⁠, ⁠$metadata⁠, ⁠$html⁠, ⁠$file⁠, ⁠$results⁠, and ⁠$call⁠.

See Also

vars(), tab(), tabexport(), usedf()

Other R4VN tables: tab(), tabexport(), tabforest(), tabmeta(), tabmulti(), tabscale(), tabscore(), tabsurvey(), vars()

Examples

# -------------------------------------------------------------------------
# 1. Repeated cross-sectional continuous outcome: no optional package needed
# -------------------------------------------------------------------------
set.seed(11)
d <- data.frame(
  period = factor(rep(c("Before", "After"), each = 80),
                  levels = c("Before", "After")),
  group = factor(rep(rep(c("Control", "Intervention"), each = 40), 2)),
  age = rnorm(160, 45, 10)
)
d$score <- 50 + 2 * (d$period == "After") +
  5 * (d$group == "Intervention") +
  4 * (d$period == "After" & d$group == "Intervention") +
  0.15 * d$age + rnorm(160, 0, 8)

z <- tablong(
  d, vars = vars(c.score), time = period, by = group,
  adjusted = vars(c.age), show = FALSE
)
z$data
z$tests_table
z$contrasts_table


# -----------------------------------------------------------------------
# 2. The same analysis with interpretation and a publication profile plot
# -----------------------------------------------------------------------
z2 <- tablong(
  d, vars = vars(c.score), time = period, by = group,
  adjusted = vars(c.age), interpretation = TRUE,
  plot = TRUE, show = FALSE
)
z2$interpretation
z2$diagnostics
plot(z2)

# -----------------------------------------------------------------------
# 3. No comparison group: change over time only
# -----------------------------------------------------------------------
z_time <- tablong(
  d, vars = vars(c.score), time = period,
  adjusted = vars(c.age), show = FALSE
)
z_time$tests_table

# -----------------------------------------------------------------------
# 4. Change the reference time by value or by bN. prefix
# -----------------------------------------------------------------------
z_ref1 <- tablong(d, vars = vars(c.score), time = period,
                  by = group, ref = "After", show = FALSE)
z_ref2 <- tablong(d, vars = vars(c.score), time = b2.period,
                  by = group, show = FALSE)

# -----------------------------------------------------------------------
# 5. Median/IQR or full descriptive display, while inference remains a
#    mean model
# -----------------------------------------------------------------------
z_median <- tablong(d, vars = vars(q.score), time = period,
                    by = group, show = FALSE)
z_full <- tablong(d, vars = vars(f.score), time = period,
                  by = group, missing = TRUE, show = FALSE)

# -----------------------------------------------------------------------
# 6. Long repeated data: linear mixed model
# -----------------------------------------------------------------------
  set.seed(12)
  n_subject <- 60
  dl <- expand.grid(
    id = seq_len(n_subject),
    visit = factor(c("Baseline", "Month 3", "Month 6"),
                   levels = c("Baseline", "Month 3", "Month 6"))
  )
  dl <- dl[order(dl$id, dl$visit), ]
  trt <- factor(sample(c("Control", "Intervention"), n_subject, TRUE),
                levels = c("Control", "Intervention"))
  dl$treatment <- rep(trt, each = 3)
  u <- rnorm(n_subject, 0, 6)
  dl$sbp <- 140 + u[dl$id] - 3 * (dl$visit == "Month 3") -
    5 * (dl$visit == "Month 6") -
    4 * (dl$treatment == "Intervention" & dl$visit == "Month 6") +
    rnorm(nrow(dl), 0, 5)

  mixed <- tablong(
    dl, vars = vars(c.sbp), time = visit, id = id,
    by = treatment, diagnostics = TRUE, show = FALSE
  )
  mixed$models[[1]]
  mixed$diagnostics

  # ---------------------------------------------------------------------
  # 7. Wide repeated data: internally converted to long format
  # ---------------------------------------------------------------------
  dw <- data.frame(
    id = seq_len(n_subject), treatment = trt,
    sbp0 = rnorm(n_subject, 140, 10)
  )
  dw$sbp3 <- dw$sbp0 - 3 + rnorm(n_subject, 0, 4)
  dw$sbp6 <- dw$sbp0 - 5 - 4 * (dw$treatment == "Intervention") +
    rnorm(n_subject, 0, 4)
  wide <- tablong(
    dw, vars = vars(c.sbp0, c.sbp3, c.sbp6),
    time = c("Baseline", "Month 3", "Month 6"),
    id = id, by = treatment, show = FALSE
  )
  wide$input_format
  head(wide$long_data)

  # ---------------------------------------------------------------------
  # 8. Binary repeated outcome: marginal logistic model with clustered SE and OR
  # ---------------------------------------------------------------------
  p <- plogis(-1 + 0.4 * (dl$visit == "Month 6") +
                0.6 * (dl$treatment == "Intervention"))
  dl$controlled <- factor(rbinom(nrow(dl), 1, p),
                          levels = c(0, 1), labels = c("No", "Yes"))
  binary_or <- tablong(
    dl, vars = vars(controlled), time = visit, id = id,
    by = treatment, event = "Yes", show = FALSE
  )
  binary_or$contrasts_table

  # ---------------------------------------------------------------------
  # 9. Continuous time and random slope
  # ---------------------------------------------------------------------
  ds <- expand.grid(id = seq_len(50), month = c(0, 3, 6, 12))
  ds <- ds[order(ds$id, ds$month), ]
  ds$group <- factor(rep(sample(c("Control", "Intervention"), 50, TRUE), each = 4))
  b0 <- rnorm(50, 0, 5)
  b1 <- rnorm(50, 0, 0.15)
  ds$score <- 50 + b0[ds$id] + (-0.3 + b1[ds$id]) * ds$month -
    0.25 * ds$month * (ds$group == "Intervention") + rnorm(nrow(ds), 0, 3)
  slope_fit <- tablong(
    ds, vars = vars(c.score), time = c.month, id = id,
    by = group, slope = TRUE, show = FALSE
  )
  slope_fit$contrasts_table

  # ---------------------------------------------------------------------
  # 10. Repeated count outcome with person-time offset -> IRR
  # ---------------------------------------------------------------------
  dc <- dl
  dc$person_time <- runif(nrow(dc), 0.8, 1.2)
  rate <- exp(0.2 + 0.2 * (dc$visit == "Month 6") -
                0.3 * (dc$treatment == "Intervention" & dc$visit == "Month 6"))
  dc$events <- rpois(nrow(dc), rate * dc$person_time)
  count_fit <- tablong(
    dc, vars = vars(c.events), time = visit, id = id,
    by = treatment, count = TRUE, exposure = person_time,
    show = FALSE
  )
  count_fit$contrasts_table

# -----------------------------------------------------------------------
# 11. RR/PR without extra packages; optional AR(1) GEE
# -----------------------------------------------------------------------
set.seed(13)
n_subject <- 70
dg <- expand.grid(
  id = seq_len(n_subject),
  visit = factor(c("Baseline", "Month 6"),
                 levels = c("Baseline", "Month 6"))
)
dg <- dg[order(dg$id, dg$visit), ]
dg$treatment <- factor(
  rep(sample(c("Control", "Intervention"), n_subject, TRUE), each = 2),
  levels = c("Control", "Intervention")
)
p <- plogis(-1 + 0.3 * (dg$visit == "Month 6") +
              0.4 * (dg$treatment == "Intervention"))
dg$controlled <- factor(rbinom(nrow(dg), 1, p),
                        levels = c(0, 1), labels = c("No", "Yes"))

fit_rr <- tablong(
  dg, vars = vars(controlled), time = visit, id = id,
  by = treatment, event = "Yes", rr = TRUE, show = FALSE
)
fit_pr <- tablong(
  dg, vars = vars(controlled), time = visit, id = id,
  by = treatment, event = "Yes", pr = TRUE,
  show = FALSE
)
fit_rr$contrasts_table
fit_pr$contrasts_table

# AR(1) is advanced and uses geepack only when explicitly requested.
if (requireNamespace("geepack", quietly = TRUE)) {
  fit_pr_ar1 <- tablong(
    dg, vars = vars(controlled), time = visit, id = id,
    by = treatment, event = "Yes", pr = TRUE, gee = TRUE, ar1 = TRUE,
    show = FALSE
  )
}

# -----------------------------------------------------------------------
# 12. Three or more groups: keep main table compact, request pairwise tests
# -----------------------------------------------------------------------
set.seed(14)
dm <- data.frame(
  period = factor(rep(c("Baseline", "Follow-up"), each = 90),
                  levels = c("Baseline", "Follow-up")),
  arm = factor(rep(rep(c("A", "B", "C"), each = 30), 2))
)
dm$score <- rnorm(nrow(dm), 50 + 2 * (dm$period == "Follow-up") +
                    2 * (dm$arm == "B") + 4 * (dm$arm == "C"), 7)
multi_arm <- tablong(
  dm, vars = vars(c.score), time = period, by = arm,
  pairwise = TRUE, adjust = "holm", show = FALSE
)
multi_arm$tests_table
multi_arm$contrasts_table

# -----------------------------------------------------------------------
# 13. Several outcomes in one long-data analysis
# -----------------------------------------------------------------------
d$positive <- factor(
  rbinom(nrow(d), 1, plogis(-1 + 0.5 * (d$period == "After"))),
  levels = c(0, 1), labels = c("No", "Yes")
)
multi_outcome <- tablong(
  d, vars = vars(c.score, positive), time = period,
  by = group, event = "Yes", show = FALSE
)
multi_outcome$data
multi_outcome$descriptive

# -----------------------------------------------------------------------
# 14. Export remains compatible with ordinary R4VN table workflows
# -----------------------------------------------------------------------
export_data <- tabexport(z)
head(export_data)

# -----------------------------------------------------------------------
# 15. Missing outcomes/attrition: inspect counts before publication
# -----------------------------------------------------------------------
d_missing <- d
d_missing$score[c(2, 7, 21, 100)] <- NA
miss_fit <- tablong(
  d_missing, vars = vars(c.score), time = period, by = group,
  missing = TRUE, diagnostics = TRUE, show = FALSE
)
miss_fit$descriptive
miss_fit$diagnostics

# -----------------------------------------------------------------------
# 16. Several binary outcomes can use a named event vector
# -----------------------------------------------------------------------
d$admitted <- factor(
  rbinom(nrow(d), 1, plogis(-1.4 + 0.4 * (d$period == "After"))),
  levels = c(0, 1), labels = c("No", "Yes")
)
binary_set <- tablong(
  d, vars = vars(positive, admitted), time = period, by = group,
  event = c(positive = "Yes", admitted = "Yes"), show = FALSE
)
binary_set$tests_table

# -----------------------------------------------------------------------
# 17. Continuous Gaussian GEE with AR(1) working correlation
# -----------------------------------------------------------------------
if (requireNamespace("geepack", quietly = TRUE)) {
  set.seed(17)
  dg2 <- expand.grid(id = seq_len(60), month = c(0, 3, 6, 12))
  dg2 <- dg2[order(dg2$id, dg2$month), ]
  dg2$group <- factor(rep(sample(c("Control", "Intervention"), 60, TRUE), each = 4))
  dg2$score <- 55 - 0.2 * dg2$month -
    0.15 * dg2$month * (dg2$group == "Intervention") + rnorm(nrow(dg2), 0, 5)
  gee_cont <- tablong(
    dg2, vars = vars(c.score), time = c.month, id = id, by = group,
    gee = TRUE, ar1 = TRUE, show = FALSE
  )
  gee_cont$contrasts_table
}

# -----------------------------------------------------------------------
# 18. Wide count data can supply one person-time variable per time point
# -----------------------------------------------------------------------
set.seed(18)
nw <- 50
wc <- data.frame(
  id = seq_len(nw),
  group = factor(sample(c("Control", "Intervention"), nw, TRUE)),
  pt0 = runif(nw, 0.8, 1.2),
  pt6 = runif(nw, 0.8, 1.2)
)
wc$event0 <- rpois(nw, 1.2 * wc$pt0)
wc$event6 <- rpois(nw,
  exp(log(1.2) - 0.25 * (wc$group == "Intervention")) * wc$pt6)
wide_count <- tablong(
  wc, vars = vars(c.event0, c.event6),
  time = c("Baseline", "Month 6"), id = id, by = group,
  count = TRUE, exposure = vars(pt0, pt6), show = FALSE
)
wide_count$contrasts_table

# -----------------------------------------------------------------------
# 19. Replot an existing result without refitting the statistical model
# -----------------------------------------------------------------------
plot(z, ci = FALSE, title = "Observed longitudinal profile",
     base_size = 12, legend_position = "right")


R4VN documentation built on Sept. 30, 2026, 5:13 p.m.