| tabsurvey | R Documentation |
Performs R4VN-style descriptive analysis, hypothesis testing, and effect
estimation for complex survey data. The interface deliberately mirrors
tab() while adding survey weights, strata, clusters, replicate
weights, domain analysis, design-based standard errors, weighted and
unweighted results, and optional population totals.
tabsurvey(
data = NULL,
vars = NULL,
by = NULL,
design = NULL,
weight = NULL,
strata = NULL,
cluster = NULL,
fpc = NULL,
repweights = NULL,
rep_type = NULL,
weightscale = c("relative", "population"),
nest = TRUE,
subpop = NULL,
result = c("weighted", "unweighted", "both"),
bothstyle = c("columns", "rows"),
statcols = c("separate", "compact"),
rawn = TRUE,
digit = 1,
p_digit = 3,
effect_digit = 2,
level = 0.95,
missing = c("ifany", "no", "always"),
row = FALSE,
col = TRUE,
cell = FALSE,
overall = c("first", "last", "none"),
descriptive = TRUE,
rvrow = NULL,
rvcol = FALSE,
test = TRUE,
pvalue = TRUE,
survey_test = c("F", "Chisq", "Wald", "adjWald"),
or = FALSE,
rr = FALSE,
pr = FALSE,
event = NULL,
adjusted = NULL,
multi = NULL,
effect_ref = NULL,
ci = TRUE,
cimethod = c("logit", "likelihood", "beta", "mean", "asin", "xlogit"),
quantile_method = c("mean", "beta", "xlogit", "asin", "score", "quantile"),
se = FALSE,
deff = FALSE,
cv = FALSE,
population = FALSE,
lonely = NULL,
bold_p = TRUE,
p_bold = 0.05,
test_note = TRUE,
template = c("journal", "clean", "minimal"),
append = NULL,
file = NULL,
raw = FALSE,
name = FALSE,
title = NULL,
report = c("auto", "brief", "full", "custom"),
interpretation = FALSE,
show = TRUE
)
data |
Optional data frame. Normally omitted when a stored
|
vars |
Variables to summarize, created with |
by |
Optional outcome/grouping variable. An unprefixed variable is
treated as categorical. Use |
design |
Survey design. May be an |
weight, strata, cluster, fpc |
Direct design arguments for one-off
analyses. These are alternatives to |
repweights |
Optional replicate weights for a one-off design. |
rep_type |
Replicate design type when |
weightscale |
|
nest |
Logical for a one-off multistage design. |
subpop |
Optional logical domain/subpopulation expression, for example
|
result |
Which analysis system to show:
|
bothstyle |
When |
statcols |
Presentation of descriptive statistics and effect estimates.
|
rawn |
Include the actual unweighted sample n in descriptive cells.
This is especially important beside weighted estimates and also keeps n
visible for unweighted continuous summaries. The default is |
digit |
Decimal places for descriptive estimates. |
p_digit |
Decimal places for p-values. |
effect_digit |
Decimal places for OR, PR, RR, and beta estimates. |
level |
Confidence level. The default is 0.95. |
missing |
|
row, col, cell |
Percentage denominator for categorical variables when
|
overall |
Position of the overall descriptive column:
|
descriptive |
Logical. Include descriptive statistics. |
rvrow |
Optional categorical row reversal, matching |
rvcol |
Logical. Reverse the displayed levels of a categorical
|
test |
Logical. Include omnibus/group-comparison tests. |
pvalue |
Logical. Include coefficient-level p-values beside effect estimates. |
survey_test |
Statistic for categorical design-adjusted association
tests passed to |
or |
Logical. For a binary categorical outcome, estimate odds ratios using logistic regression. |
rr |
Logical. For a binary outcome, estimate risk/prevalence ratios with a log-link modified Poisson model. In cross-sectional surveys this is interpreted as a prevalence ratio. |
pr |
Logical. Estimate prevalence ratios with a log-link modified
Poisson model. Weighted models use |
event |
Event level for a binary categorical outcome. By default the last observed outcome level is the event. |
adjusted |
Optional adjustment set. Supply |
multi |
Optional multivariable set. Supply |
effect_ref |
Optional backward-compatible explicit reference mapping
for crude and separately adjusted categorical effects, for example
|
ci |
Logical. Show confidence intervals at the selected |
cimethod |
Confidence-interval method for weighted proportions:
|
quantile_method |
Interval method used by
|
se |
Logical. Add a separate standard-error column for descriptive estimates. When unweighted results are requested, their conventional SE is also reported where defined. |
deff |
Logical. Add a separate with-replacement design-effect column for weighted statistics where the underlying survey statistic supports it. |
cv |
Logical. Add a separate coefficient-of-variation/relative-SE column where defined for weighted and unweighted descriptive estimates. |
population |
Logical. Append estimated population N and its confidence interval
for categorical cells. This requires a design declared with
|
lonely |
Optional lonely-PSU rule for this analysis. If omitted, the rule stored in the design is used. |
bold_p |
Logical. Bold p-values smaller than |
p_bold |
Threshold used when |
test_note |
Logical. Add footnotes describing the tests used. |
template |
HTML style: |
append |
Optional previous R4VN table object to place before this table in the generated HTML page. |
file |
Optional HTML output path. A temporary file is used when omitted. |
raw |
Logical. Use raw variable names instead of variable labels. |
name |
Logical. When labels exist, append the raw variable name in square brackets. |
title |
Optional table title. |
report |
Reporting profile: |
interpretation |
Logical. Add a cautious deterministic interpretation
table. The default is |
show |
Logical. Open the generated HTML report in the Viewer/browser. |
Dependency-light implementation.
Beyond R4VN itself, tabsurvey() requires only the survey
package for complex-survey estimation. Publication HTML is generated with
base R; ggplot2, plotly, htmlwidgets, flextable,
and similar presentation packages are not required. tabsurvey() is
a table/inference function and does not create a plot, so it deliberately
adds no plotting dependency. R4VN functions that do create plots should
embed every requested plot directly in their Viewer/HTML report.
Weighted and unweighted are complete analysis modes.
With result = "both", R4VN computes two parallel analyses. The
unweighted side uses ordinary sample descriptions and conventional tests or
regressions. The weighted side uses the declared survey design for
descriptive estimates, standard errors, confidence intervals, Rao-Scott or
design-based tests, and survey-weighted regression. This is intentionally
more comprehensive than merely displaying a raw n beside a weighted
percentage.
Default publication display.
The default statcols = "separate" uses distinct columns for sample n,
estimate, and confidence interval instead of combining them in one long cell. Optional
SE, DEFF, CV, population totals, model effects, model confidence intervals,
and model p-values are also separate columns. The default
result = "weighted", rawn = TRUE shows the actual sample n together
with the survey-weighted estimate. For categorical variables the
weighted statistic is a percentage with a design-based confidence interval. For
c. variables the weighted mean and weighted population SD are shown,
with a design-based CI for the mean. For q. variables the weighted
median and weighted IQR are shown, with a median CI when available.
Full summaries.
A variable declared with f. produces separate mean (SD),
median (IQR), and range rows so weighted and unweighted summaries can be
compared without compressing incompatible statistics into one number.
Tests.
For categorical predictor by categorical outcome, weighted inference uses
survey::svychisq() and defaults to the second-order Rao-Scott F
correction. Weighted continuous comparisons use design-based t/Wald tests
for mean-oriented variables and survey::svyranktest() for
median/rank-oriented variables.
Regression estimates.
OR uses survey-weighted logistic regression. PR and RR use a log-link
survey-weighted quasi-Poisson model. A continuous by = c.outcome or
by = q.outcome automatically reports unstandardized beta
coefficients; the q. prefix changes the descriptive/group test but
beta remains a linear-regression coefficient, consistent with R4VN
tab() conventions.
Reference categories.
Categorical references follow vars() prefixes. For example
b2.sex makes the second observed/displayed level the model reference.
The same requested reference is used in weighted and unweighted models.
Domain analysis.
Use subpop= instead of physically deleting observations and
rebuilding a complex design. The survey domain/subset machinery keeps
the design information needed for valid variance estimation.
Population totals.
population = TRUE is intentionally blocked unless
weightscale = "population". Weighted percentages, means, tests and
regressions remain valid with normalized/relative survey weights, but their
sum must not automatically be interpreted as the represented population.
Continuous outcomes.
When by is continuous, predictor descriptions remain available and
association tests/effect columns concern the continuous outcome. Categorical
predictors are compared with t/ANOVA or rank tests as appropriate; numeric
predictors are assessed by the slope test. The effect is an unstandardized
beta coefficient with a confidence interval.
Replicate-weight designs.
Replicate weights may be defined in surveyset() or directly in
tabsurvey(). All statistics are then delegated to the corresponding
survey replicate-design methods.
Reporting profiles.
report = "auto" is the recommended default: it keeps the main table
compact and weighted, automatically includes design-based tests when a
by variable is present, and shows supporting design/test/effect
tables in the Viewer. "brief" is deliberately descriptive.
"full" adds the unweighted comparison plus SE, DEFF, and CV and
stacks weighted/unweighted results by rows to avoid excessively wide
tables. "custom" preserves the older option-by-option behavior.
Interpretation is never automatic; set interpretation = TRUE.
Invisibly returns an object of classes
r4vn_tabsurvey, r4vn_tab, and list. Important
components include:
data: flat publication-ready table, compatible with
tabexport();
html, table_html, and file: rendered table;
design: the R4VN survey design metadata;
survey_design: the underlying survey design used after
any domain restriction;
metadata: resolved R4VN variable specifications;
tests: long-form machine-friendly test results;
effects: long-form machine-friendly OR/PR/RR/beta results;
notes: table footnotes;
tables: named end-user report tables including Main, Design, Tests, Effects, Precision, and Interpretation when available;
diagnostics: survey-design and precision diagnostics;
models: fitted survey/unweighted regression models used for reported effects;
interpretation: optional deterministic interpretation table;
subpop: domain expression, when used.
For a publication or survey report, describe the sampling design and source of the final analytic weight, identify strata and PSU variables, state any domain/subpopulation restriction, and report the actual sample n together with survey-weighted estimates and design-based confidence intervals. When a hypothesis test is reported, the survey-adjusted test should normally be treated as the inferential result for a complex probability sample.
When result = "both", the unweighted analysis is useful for data
checking, transparency, and showing how weighting/design affects the
result; it does not replace the design-based inference.
Do not interpret the sum of normalized/relative weights as a
population size. Use weightscale = "population" only when the
survey documentation supports an expansion-weight interpretation.
Do not create a survey domain by deleting all observations outside
the target subgroup and rebuilding the design. Prefer
subpop = ....
Do not assume one weight is correct for every variable in a public
survey. When different analytic components require different weights,
create multiple named designs with surveyset().
Do not silently treat propensity-score IPTW, frequency weights, or analytic regression weights as sampling/design weights.
Lumley T. Complex Surveys: A Guide to Analysis Using R. Wiley; 2010.
Lumley T. Analysis of complex survey samples. Journal of Statistical Software. 2004;9(1):1-19.
surveyset, tab, vars,
tabexport
Other R4VN survey:
surveyset()
Other R4VN tables:
tab(),
tabexport(),
tabforest(),
tablong(),
tabmeta(),
tabmulti(),
tabscale(),
tabscore(),
vars()
# Reproducible complex-survey data used throughout the examples.
set.seed(2026)
d <- expand.grid(
person = 1:2, household = 1:5, psu = 1:6, strata = 1:4,
KEEP.OUT.ATTRS = FALSE
)
n <- nrow(d)
d$sex <- factor(sample(c("Female", "Male"), n, TRUE),
levels = c("Female", "Male"))
d$age <- pmin(85, pmax(18, round(rnorm(n, 46, 14))))
d$bmi <- round(rnorm(n, 23.5, 3.4), 1)
d$income <- round(exp(rnorm(n, log(8), .5)), 1)
d$smoking <- factor(sample(c("No", "Yes"), n, TRUE, c(.72, .28)),
levels = c("No", "Yes"))
d$education <- factor(
sample(c("Primary", "Secondary", "College+"), n, TRUE),
levels = c("Primary", "Secondary", "College+")
)
d$wt <- exp(.15 * (d$sex == "Male") + rnorm(n, 0, .3))
d$labwt <- d$wt * exp(rnorm(n, 0, .12))
d$popwt <- d$wt * 5000
d$fpc1 <- 30
d$fpc2 <- 100
lp <- -5 + .055 * d$age + .08 * (d$bmi - 23) +
.45 * (d$sex == "Male") + .55 * (d$smoking == "Yes")
d$hypertension <- factor(rbinom(n, 1, plogis(lp)),
levels = 0:1, labels = c("No", "Yes"))
d$sbp <- 82 + .72 * d$age + .85 * d$bmi +
5 * (d$sex == "Male") + rnorm(n, 0, 13)
# Declare the survey design once; later tabsurvey() calls can stay short.
usedf(d)
surveyset(weight = wt, strata = strata, cluster = psu, nest = TRUE)
# 1. Simplest weighted publication table. report="auto" is the default.
s1 <- tabsurvey(vars = vars(c.age, sex, c.bmi, smoking), show = FALSE)
s1$tables$Main
s1$tables$Design
# 2. R4VN continuous prefixes: c.=mean, q.=median, f.=full summary.
s2 <- tabsurvey(vars = vars(c.age, q.income, f.bmi, sex), show = FALSE)
# 3. Table by a binary outcome; design-based tests are automatic.
s3 <- tabsurvey(
vars = vars(c.age, sex, c.bmi, smoking, education),
by = hypertension, show = FALSE
)
s3$tables$Tests
# 4. Compare complete unweighted and weighted analyses side by side.
s4 <- tabsurvey(
vars = vars(c.age, sex, q.income, c.bmi, smoking),
by = hypertension, result = "both", bothstyle = "columns",
show = FALSE
)
# 5. Full profile: both analyses stacked by rows plus SE, DEFF, and CV.
s5 <- tabsurvey(
vars = vars(c.age, sex, c.bmi, smoking),
by = hypertension, report = "full", show = FALSE
)
s5$tables$Precision
# 6. Brief profile: weighted descriptive summary only unless overridden.
s6 <- tabsurvey(
vars = vars(c.age, sex, q.income, c.bmi),
report = "brief", show = FALSE
)
# 7. Row or cell percentages instead of the default column percentages.
s7_row <- tabsurvey(
vars = vars(sex, smoking, education), by = hypertension,
row = TRUE, col = FALSE, cell = FALSE, show = FALSE
)
s7_cell <- tabsurvey(
vars = vars(sex, smoking, education), by = hypertension,
row = FALSE, col = FALSE, cell = TRUE, show = FALSE
)
# 8. Crude survey-weighted odds ratios in the same publication table.
s8 <- tabsurvey(
vars = vars(c.age, b2.sex, c.bmi, b2.smoking, education),
by = hypertension, or = TRUE, event = "Yes", show = FALSE
)
s8$tables$Effects
# 9. Separately adjusted OR for every focal predictor.
s9 <- tabsurvey(
vars = vars(c.age, b2.sex, c.bmi, b2.smoking),
by = hypertension, or = TRUE, event = "Yes",
adjusted = vars(c.age, b2.sex), show = FALSE
)
# 10. One common multivariable model containing all requested predictors.
s10 <- tabsurvey(
vars = vars(c.age, b2.sex, c.bmi, b2.smoking, education),
by = hypertension, or = TRUE, event = "Yes",
multi = TRUE, show = FALSE
)
s10$models
# 11. Prevalence ratio via survey-weighted modified Poisson regression.
s11 <- tabsurvey(
vars = vars(c.age, b2.sex, c.bmi, b2.smoking),
by = hypertension, pr = TRUE, event = "Yes",
multi = TRUE, show = FALSE
)
# 12. Explicit named reference levels; b2./b3. are also supported.
s12 <- tabsurvey(
vars = vars(sex, smoking, education, c.age),
by = hypertension, or = TRUE, event = "Yes",
effect_ref = list(sex = "Male", smoking = "Yes",
education = "Secondary"),
show = FALSE
)
# 13. Continuous outcome: unstandardized beta is reported automatically.
s13 <- tabsurvey(
vars = vars(c.age, b2.sex, c.bmi, b2.smoking),
by = c.sbp, multi = TRUE, result = "both", show = FALSE
)
# 14. q. continuous outcome requests rank-oriented group tests; effect is beta.
s14 <- tabsurvey(
vars = vars(b2.sex, b2.smoking, education),
by = q.sbp, result = "both", show = FALSE
)
# 15. Correct domain/subpopulation analysis; do not rebuild a reduced design.
s15 <- tabsurvey(
vars = vars(c.age, sex, c.bmi, smoking),
subpop = age >= 60 & sex == "Female", show = FALSE
)
s15$diagnostics$domain
# 16. Missing rows can be shown if present, always, or never.
d$smoking[1:4] <- NA
surveyset(d, name = "missing_demo", weight = wt, strata = strata,
cluster = psu)
s16 <- tabsurvey(
vars = vars(smoking, sex), design = "missing_demo",
missing = "ifany", show = FALSE
)
# 17. Confidence level is fully dynamic, including the displayed CI label.
s17 <- tabsurvey(
vars = vars(c.age, sex, c.bmi), by = hypertension,
or = TRUE, event = "Yes", level = .90, show = FALSE
)
names(s17$data) # contains "90% CI"
# 18. Request SE, design effect, and CV explicitly in a custom report.
s18 <- tabsurvey(
vars = vars(c.age, sex, c.bmi, smoking),
report = "custom", se = TRUE, deff = TRUE, cv = TRUE,
show = FALSE
)
s18$tables$Precision
# 19. Population totals require declared expansion/population weights.
surveyset(d, name = "population", weight = popwt, strata = strata,
cluster = psu, weightscale = "population", active = FALSE)
s19 <- tabsurvey(
vars = vars(sex, education, hypertension), design = "population",
population = TRUE, show = FALSE
)
# 20. One-off design: no prior surveyset() call is required.
s20 <- tabsurvey(
d, vars = vars(c.age, sex, c.bmi, hypertension),
weight = wt, strata = strata, cluster = psu, show = FALSE
)
# 21. Multiple named weight systems can coexist.
surveyset(d, name = "laboratory", weight = labwt,
strata = strata, cluster = psu, active = FALSE)
s21 <- tabsurvey(
vars = vars(c.age, sex, c.bmi), design = "laboratory", show = FALSE
)
# 22. Multistage clusters and finite-population corrections.
surveyset(
d, name = "multistage", weight = wt, strata = strata,
cluster = vars(psu, household), fpc = vars(fpc1, fpc2),
nest = TRUE, active = FALSE
)
s22 <- tabsurvey(
vars = vars(c.age, sex, c.bmi), design = "multistage", show = FALSE
)
# 23. Compact legacy cells and display-order controls.
s23 <- tabsurvey(
vars = vars(sex, smoking), by = hypertension,
statcols = "compact", rvrow = TRUE, rvcol = TRUE,
report = "custom", show = FALSE
)
# 24. Interpretation is opt-in and remains separate from statistical output.
s24 <- tabsurvey(
vars = vars(c.age, b2.sex, c.bmi, b2.smoking),
by = hypertension, pr = TRUE, event = "Yes", multi = TRUE,
report = "full", interpretation = TRUE, show = FALSE
)
s24$tables$Interpretation
# 25. Consistent result contract for custom reporting and downstream code.
names(s24$tables)
s24$descriptive
s24$tests
s24$effects
s24$diagnostics
s24$models
s24$interpretation
# 26. tabsurvey objects inherit from r4vn_tab and export with tabexport().
h <- tabexport(
s3, s8, s11, s24, export = "html",
file = tempfile("survey_report_"), quiet = TRUE
)
unlink(h$files)
# Additional syntax catalogue. These examples are intentionally not run by
# automatic checks, but are kept in ?tabsurvey for copy/paste use.
# 27. Risk ratio using the same modified-Poisson engine.
s27 <- tabsurvey(
vars = vars(c.age, b2.sex, c.bmi, b2.smoking),
by = hypertension, rr = TRUE, event = "Yes", multi = TRUE
)
# 28. Alternative CI methods for proportions and weighted quantiles.
s28_prop <- tabsurvey(
vars = vars(sex, smoking, hypertension), cimethod = "beta"
)
s28_quantile <- tabsurvey(
vars = vars(q.income, q.bmi), quantile_method = "beta"
)
# 29. Choose another design-adjusted categorical test.
s29 <- tabsurvey(
vars = vars(sex, smoking, education), by = hypertension,
survey_test = "Wald"
)
# 30. Display controls: no CI, no raw n, two decimals, Overall last.
s30 <- tabsurvey(
vars = vars(c.age, sex, c.bmi, smoking), by = hypertension,
ci = FALSE, rawn = FALSE, digit = 2, overall = "last"
)
# 31. Variable-name and HTML presentation controls.
s31_raw <- tabsurvey(
vars = vars(c.age, sex, c.bmi), raw = TRUE,
template = "clean", title = "Raw variable names"
)
s31_name <- tabsurvey(
vars = vars(c.age, sex, c.bmi), name = TRUE,
template = "minimal", title = "Labels plus names"
)
# 32. Inference/model-only table with descriptive cells suppressed.
s32 <- tabsurvey(
vars = vars(c.age, b2.sex, c.bmi, b2.smoking),
by = hypertension, or = TRUE, event = "Yes", multi = TRUE,
descriptive = FALSE, report = "custom"
)
# 33. Explicitly suppress tests, coefficient p-values, and test notes.
s33 <- tabsurvey(
vars = vars(c.age, b2.sex, c.bmi), by = hypertension,
or = TRUE, event = "Yes", test = FALSE, pvalue = FALSE,
test_note = FALSE, bold_p = FALSE, report = "custom"
)
# 34. Append two R4VN survey tables into one HTML page.
a34 <- tabsurvey(vars = vars(c.age, sex), show = FALSE)
f34 <- tempfile(fileext = ".html")
b34 <- tabsurvey(
vars = vars(c.bmi, smoking), append = a34,
file = f34, show = FALSE
)
unlink(f34)
# 35. A survey-package replicate design can be passed directly.
base35 <- survey::svydesign(
ids = ~psu, strata = ~strata, weights = ~wt, data = d, nest = TRUE
)
rep35 <- survey::as.svrepdesign(base35, type = "bootstrap", replicates = 40)
s35 <- tabsurvey(
vars = vars(c.age, sex, c.bmi, hypertension),
design = rep35, report = "full"
)
# 36. Override the lonely-PSU rule for one analysis only.
s36 <- tabsurvey(
vars = vars(c.age, sex, c.bmi), lonely = "average"
)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.