| tabscore | R Documentation |
tabscore() converts a multivariable prediction model into a complete,
publication-ready scorecard. The function is designed for the full workflow,
not merely for rounding regression coefficients. In one call it can:
develop or accept a final prediction model;
optionally select predictors from a candidate set;
create score-ready categories for continuous predictors;
derive an exact/model score and a simpler clinical integer score;
produce a complete Predictor–Category–Point table and theoretical total score range;
map every possible total score to predicted risk (or expected count/rate);
select clinically/statistically useful score cutoffs;
report diagnostic/prognostic properties at the selected cutoff;
compare the original model, the model score and the clinical score;
perform bootstrap internal validation of the entire development pipeline;
optionally evaluate an external validation data set; and
retain plot-ready data and prediction methods for deployment in R4VN Studio.
Version 1 supports logistic regression, Cox proportional hazards regression, and Poisson regression. Logistic and Cox models receive the most complete discrimination/cutoff workflow. A Poisson model may be used for count/rate scores; when its outcome is binary, ROC/cutoff summaries are also available.
tabscore(
outcome,
predictors = NULL,
data = NULL,
family = c("auto", "logistic", "cox", "poisson"),
event = NULL,
time = NULL,
times = NULL,
cutoff_time = NULL,
select = c("none", "full", "backward", "forward", "purposeful", "lasso"),
force = NULL,
exclude = NULL,
entry = 0.25,
stay = 0.1,
confound = 0.15,
cuts = NULL,
continuous = c("easy", "auto", "quantile", "keep"),
bins = 4L,
points = "auto",
pdo = 20,
maxscore = "auto",
simplify = TRUE,
tolerance = 0.01,
riskonly = TRUE,
scoreref = "lowest",
cutoff = c("iu", "youden", "risk", "prevalence", "sens", "spec", "cost", "manual",
"refprob", "none"),
riskcut = NULL,
cutoff_value = NULL,
sens = NULL,
spec = NULL,
cost_fp = 1,
cost_fn = 1,
refprob = NULL,
refcut = NULL,
validate = c("bootstrap", "none"),
bootstrap = 500L,
validation = NULL,
risktable = TRUE,
compare = TRUE,
calibration = TRUE,
decision = TRUE,
plot = TRUE,
show = TRUE,
console = FALSE,
seed = NULL,
ai = FALSE
)
outcome |
Outcome variable. It can be an unquoted variable name, a
character variable name, an already fitted |
predictors |
Candidate predictors. Accepts a character vector, unquoted
variables inside |
data |
Data frame. When omitted, |
family |
Model family: |
event |
Event level for a binary outcome/status. For 0/1 outcomes the
default is 1. For a two-level factor the default is its second factor level;
for character outcomes it is the second observed non-missing value. Set this explicitly whenever the event
direction matters, e.g. |
time |
Cox follow-up-time variable, supplied as an unquoted or character variable name. This is the observed time, not the prediction horizon. |
times |
Prediction horizons for Cox score-to-risk tables, e.g.
|
cutoff_time |
Time horizon at which a Cox cutoff is evaluated. Defaults
to the largest value in |
select |
Predictor-selection strategy. |
force |
Predictors that must remain in model-selection procedures. Accepts
character names, |
exclude |
Candidate predictors to remove before model development. Accepts
character names, |
entry |
Univariable screening p-value for purposeful selection; default 0.25. |
stay |
Multivariable retention/re-entry p-value for purposeful selection; default 0.10. |
confound |
Relative coefficient-change threshold used to retain a variable as a confounder during purposeful selection; default 0.15 (15 percent). |
cuts |
Optional named list of user-defined cut points for continuous
predictors, e.g. |
continuous |
Handling of continuous predictors when the clinical scorecard
is built. |
bins |
Desired number of categories when automatic continuous-variable categorization is used. Default 4. |
points |
Point-construction method. |
pdo |
Points to double the effect on the model's log scale. For logistic
regression this is Points to Double the Odds: |
maxscore |
Maximum preferred theoretical clinical score. |
simplify |
Logical. With |
tolerance |
Maximum tolerated decrease in the primary discrimination metric when simplifying to the clinical score. For binary models this is AUC; for Cox it is C-index. The same threshold is also used to flag excessive loss caused by automatic predictor categorization. Default 0.01. If no compact score meets the tolerance, the best-performing candidate is retained and a warning is stored. |
riskonly |
Logical. Default TRUE. Build the bedside score as a pure add-only risk score: each predictor is re-referenced to its lowest modeled risk category, all scoring effect ratios are at least 1, and all clinical points are non-negative. For example, if Male is the regression reference and Female has OR=0.50, the score representation becomes Female=0 points (score reference) and Male has risk-oriented OR=2.00 with positive points. This reparameterization does not change the fitted model, subject ranking, or predicted risks. Set FALSE only when a signed score with negative protective points is specifically desired. |
scoreref |
Scoring reference. Default |
cutoff |
Selected cutoff principle: |
riskcut |
One or more clinically meaningful probability thresholds. With
|
cutoff_value |
User-specified integer score threshold for |
sens |
Minimum desired sensitivity for |
spec |
Minimum desired specificity for |
cost_fp |
Relative cost assigned to a false positive. Default 1. |
cost_fn |
Relative cost assigned to a false negative. Example: |
refprob |
Optional reference predicted probability for binary-outcome
workflows: a numeric vector or a probability-variable name in |
refcut |
Probability threshold applied to |
validate |
Internal validation: |
bootstrap |
Number of bootstrap resamples. Default 500. Use 1000 or more for a final analysis when feasible; small values are for code testing only. |
validation |
Optional external validation data frame. The frozen final scorecard is applied without re-estimating cuts, points, calibration or the chosen score cutoff. Binary validation reports AUC, Brier, calibration and fixed-cutoff performance; Cox validation reports C-index, time-specific IPCW Brier and fixed-cutoff IPCW performance; count Poisson reports prediction error. |
risktable |
Logical; create score-to-risk/score-to-expected-value table. |
compare |
Logical; compare original model, model score and clinical score. |
calibration |
Logical; prepare calibration data for plots. |
decision |
Logical; prepare decision-curve net-benefit data for binary outcomes. |
plot |
Logical. Default TRUE. Prepare all graphics supported by the chosen model family, embed every available graph directly in the HTML Viewer, and draw them into the interactive R/RStudio Plot history. Set FALSE when only tables are wanted. Plotting uses base R and does not add a required package. |
show |
Logical. If TRUE, write a self-contained publication-oriented HTML result and open it in the RStudio Viewer (or the default browser). |
console |
Logical. If TRUE, print a concise console summary. |
seed |
Optional random seed used for LASSO, bootstrap fallback and validation. The default |
ai |
Logical or endpoint name. If R4VN |
When tabscore() develops a model from candidate predictors, it uses one
complete-case development sample across the outcome/status, Cox time (when
applicable), and all candidate predictors supplied before model selection.
This keeps candidate models comparable but can reduce sample size when many
predictors have missing values. Perform the intended imputation or missing-data
strategy before tabscore() when complete-case analysis is inappropriate.
tabscore() understands the R4VN variable culture rather than merely stripping
prefixes. For example, vars(c.age, b2.sex, smoking) fits age continuously,
treats sex as categorical with its second factor level as model reference, and
treats smoking as categorical with its first factor level as reference. This
affects the fitted model table and model-selection calculations. Point assignment
itself is then shifted within each predictor so the lowest-risk category receives
zero automatic points; consequently the zero-point category does not have to be
the regression reference category. The final prediction-model table preserves
the original statistical reference and may therefore legitimately show OR/HR/RR
below 1. The separate risk_orientation table shows the scoring contrast after
re-referencing; with riskonly=TRUE, every displayed scoring ratio is >=1 and
the clinical score contains only zero or positive points. With ordinary
c(age, sex) syntax, data type is inferred from the columns instead of imposing
R4VN categorical declarations.
tabscore() distinguishes three objects. The original model is the selected
or fixed model using the original predictor representation. The model score
is a monotone point transformation of the score-ready model linear predictor.
With pdo=20, a 20-point increase doubles odds (logistic), hazard (Cox), or
modeled rate (Poisson). The clinical score uses small integer points and is
the score shown in the clean Predictor–Category–Point table.
Intercepts are never artificially divided among predictors; they remain in the risk mapping. Within each predictor, the lowest modeled contribution is shifted to zero before automatic non-negative points are assigned. Thus a zero-point category need not be the regression reference if another category has lower risk.
With riskonly=TRUE, categorical contributions are transformed predictor by
predictor as beta_score = beta_category - min(beta_categories). Therefore
exp(beta_score) >= 1. For a binary predictor with an original protective
contrast OR=0.50, reversing the scoring contrast gives 1/0.50=2.00 for the
higher-risk category. For multi-level predictors the same minimum-risk rebasing
is used; the procedure is not abs(beta), which can distort category ordering.
Cox HR and Poisson RR/IRR are handled identically on their log-effect scales.
For a continuous coefficient retained without categories, a negative coefficient
is described technically as risk per unit decrease; a complete bedside integer
score still requires explicit/automatic categories. The transformation changes
only the score origin/reference: the original fitted model and its absolute-risk
predictions remain untouched.
A base glm/coxph model or an R4VN regression result containing raw$model
may be passed as outcome. In that workflow the fitted predictor set is fixed,
so select, force and exclude are not used. This version deliberately
rejects interactions, transformed/spline terms, no-intercept models, non-unit
analysis weights, and non-zero offsets/exposures rather than silently changing
the fitted model during score simplification. Represent required transformed
predictors as explicit columns and refit before calling tabscore(), or provide
a manual point system.
Automatic scorecard categorization is a simplification step rather than part of
the original continuous model. continuous="easy" uses empirical quantiles and
snaps thresholds toward simple numbers. Prefer clinically established cuts=
when available. Because automatic cuts are data-driven, bootstrap validation
repeats the cut-selection step inside each resample.
The displayed range is calculated from all category combinations implied by the scorecard, not from the smallest/largest observed subject score. This keeps the bedside score stable in new data.
For logistic models, the clinical score is recalibrated by
logit(P)=a+b*Score; every possible total score receives a predicted probability
and 95 percent CI. This recalibration is important because categorization and
integer rounding mean the clinical score is no longer exactly the original LP.
Cox scorecards use a one-predictor Cox calibration model to provide risk at each
times= horizon. Poisson count scores provide expected count/rate and CI.
youden: maximize sensitivity + specificity - 1.
iu: minimize abs(sensitivity-AUC)+abs(specificity-AUC).
risk: map a clinical probability in riskcut to an integer score.
prevalence: use development event prevalence as probability threshold,
reproducing a common legacy workflow; it is not automatically clinically best.
sens: among thresholds meeting sens, maximize specificity.
spec: among thresholds meeting spec, maximize sensitivity.
cost: minimize cost_fn*FN + cost_fp*FP.
manual: use cutoff_value.
refprob: use refprob and refcut to reproduce a reference probability rule.
Cox support in this version is for standard right-censored proportional-hazards
models. Start-stop/time-dependent, multi-state, competing-risk and other complex
survival structures are not silently reduced to a simple integer score.
Cox ROC/cutoff calculations at cutoff_time use cumulative/dynamic IPCW
sensitivity and specificity with the censoring distribution estimated by
Kaplan-Meier. IPCW cutoff sensitivity, specificity, PPV, NPV, accuracy and
likelihood ratios are point estimates in the apparent cutoff table; unlike the
ordinary binary-outcome table, simple binomial confidence intervals are not
reported because censoring weights make them inappropriate. Full-pipeline
bootstrap validation supplies optimism-corrected cutoff performance and the
empirical stability interval of the selected score threshold. If an intervention
probability is known, riskcut is generally easier to interpret clinically than
a purely statistical cutoff.
A score is not accepted merely because an AUC difference is non-significant. Binary comparisons include AUC with 95 percent CI, Brier score, calibration intercept/slope, and paired AUC difference. Cox comparisons include C-index and time-specific IPCW Brier scores. Decision-curve data compare net benefit across probability thresholds.
Each bootstrap resample repeats predictor selection, automatic categorization,
point derivation, and the requested cutoff-selection rule. Apparent performance
is compared with performance when the bootstrap-derived score and cutoff are
applied to the original sample. For binary outcomes the validation table includes
optimism-corrected AUC, Brier score, calibration intercept/slope, sensitivity,
specificity, PPV, NPV, accuracy, and bootstrap cutoff stability. For Cox models
the validation table includes optimism-corrected C-index and, when a cutoff is
requested, time-dependent IPCW sensitivity, specificity, PPV, NPV and accuracy
at cutoff_time plus bootstrap cutoff stability. Mean optimism is subtracted
from the apparent final performance. This is more rigorous than bootstrapping
a fixed already-developed score.
Supply validation= to apply the frozen developed score to a separate data
set. Predictor cut points, integer points, risk mapping, and the selected score
threshold are not re-optimized in the external data. This avoids turning an
external validation into a second development exercise. After development,
predict(score_object, newdata, type="all") returns bedside score, predicted
risk/value and risk group where applicable. New factor values that were absent
from the development scorecard cannot be scored and therefore yield missing
score/risk rather than being silently assigned zero points.
With show=TRUE, the Viewer contains the publication tables and, when
plot=TRUE, every graph that is actually available for the fitted family.
Logistic/binary scorecards can show score-to-risk, ROC, calibration, decision
curve and score-distribution plots. Cox scorecards show score-to-risk across
requested horizons, time-dependent ROC at cutoff_time, and score
distribution. Poisson count scorecards show score-to-expected-value and the
observed-count distribution. The Viewer plots are rendered with base R, using
a system sans-serif font and an embedded raster image when possible; this keeps
the report self-contained and avoids a ggplot2/htmlwidgets dependency. Use
plot(result) to draw all available graphs in the RStudio Plots pane, or
plot(result, which="roc"), for example, to draw one.
The core logistic and Poisson workflows use base/recommended R only.
survival is needed only for Cox scorecards, glmnet only for
select="lasso", and pROC is optional because R4VN has a base-R fallback
for AUC calculations and paired AUC comparison. Thus users do not need to
install a large collection of packages for ordinary tabscore() analyses.
Prediction modeling is not equivalent to retaining only p<0.05 predictors. Use
subject-matter knowledge and force= for essential variables. Data-driven cuts
and cutoffs can overfit and should be validated. Interactions, spline bases,
time-varying Cox effects, competing risks and machine-learning distillation are
not silently converted into a bedside integer score in this version.
An object of class r4vn_tabscore with fitted models, score rules,
development scores/predictions, selected cutoff, theoretical range, publication
tables, technical tables, validation results and plot-ready data. Important
elements include models, scores, tables, publication_tables,
selected_predictors, selected_cutoff, score_range, riskonly,
effect_measure, plots, plot_titles and plot_data.
tables$risk_orientation explicitly compares
the score-ready model-reference effect ratio with the risk-oriented scoring ratio. The
publication_tables list contains only ready-to-export non-NULL tables and can
be passed directly to tabexport(). predict() can then score new
patients without re-estimating the scorecard.
predict.r4vn_tabscore, plot.r4vn_tabscore
Other R4VN tables:
tab(),
tabexport(),
tabforest(),
tablong(),
tabmeta(),
tabmulti(),
tabscale(),
tabsurvey(),
vars()
set.seed(2026)
n <- 220
d <- data.frame(
age = round(rnorm(n, 52, 12)),
bmi = round(rnorm(n, 24, 4), 1),
hypertension = factor(rbinom(n, 1, .30), 0:1, c("No", "Yes")),
smoking = factor(rbinom(n, 1, .25), 0:1, c("No", "Yes")),
alcohol = factor(rbinom(n, 1, .20), 0:1, c("No", "Yes"))
)
lp <- -4.2 + .04*d$age + .06*(d$bmi - 24) +
.8*(d$hypertension == "Yes") + .6*(d$smoking == "Yes")
d$event <- rbinom(n, 1, plogis(lp))
# 1. Simplest publication-ready logistic scorecard.
# show=TRUE and plot=TRUE are the user-facing defaults.
s1 <- tabscore(
event, c(age, bmi, hypertension, smoking), data=d,
validate="none", show=FALSE, plot=FALSE
)
s1$tables$scorecard
s1$tables$risk
s1$tables$comparison
# 2. R4VN variable declarations: continuous variables and chosen references.
s2 <- tabscore(
event, vars(c.age, c.bmi, b2.hypertension, b2.smoking), data=d,
validate="none", show=FALSE, plot=FALSE
)
s2$tables$model
s2$tables$risk_orientation
# 3. Clinically prespecified cut points.
s3 <- tabscore(
event, c(age, bmi, hypertension, smoking), data=d,
cuts=list(age=c(40,50,60), bmi=c(23,25,30)),
validate="none", show=FALSE, plot=FALSE
)
# 4. Apply the frozen scorecard to new patients.
newp <- data.frame(
age=c(45,68), bmi=c(24,29),
hypertension=factor(c("No","Yes"), levels=c("No","Yes")),
smoking=factor(c("Yes","No"), levels=c("No","Yes"))
)
predict(s3, newp, type="all")
# 5. Purposeful selection; force variables that must remain clinically.
s5 <- tabscore(
event, c(age,bmi,hypertension,smoking,alcohol), data=d,
select="purposeful", force=c("age","hypertension"),
validate="none", show=FALSE, plot=FALSE
)
# 6. Backward or forward AIC selection.
s6a <- tabscore(event, c(age,bmi,hypertension,smoking,alcohol), data=d,
select="backward", validate="none", show=FALSE, plot=FALSE)
s6b <- tabscore(event, c(age,bmi,hypertension,smoking,alcohol), data=d,
select="forward", validate="none", show=FALSE, plot=FALSE)
# 7. LASSO is optional and only needs glmnet for this selection method.
if (requireNamespace("glmnet", quietly=TRUE)) {
s7 <- tabscore(event, c(age,bmi,hypertension,smoking,alcohol), data=d,
select="lasso", validate="none", show=FALSE, plot=FALSE)
}
# 8. Compact score versus PDO/model-scale points.
s8a <- tabscore(event, c(age,hypertension,smoking), data=d,
cuts=list(age=c(40,50,60)), maxscore=10,
validate="none", show=FALSE, plot=FALSE)
s8b <- tabscore(event, c(age,hypertension,smoking), data=d,
cuts=list(age=c(40,50,60)), points="pdo", pdo=20,
validate="none", show=FALSE, plot=FALSE)
# 9. Completely manual bedside points; R4VN still calibrates and validates it.
s9 <- tabscore(
event, c(age,hypertension,smoking), data=d,
cuts=list(age=c(40,50,60)),
points=list(
age=c("<40"=0, "40-49"=1, "50-59"=2, ">=60"=3),
hypertension=c("No"=0,"Yes"=2),
smoking=c("No"=0,"Yes"=1)
), validate="none", show=FALSE, plot=FALSE
)
# 10. Common cutoff rules.
s10_iu <- tabscore(event, c(age,hypertension,smoking), data=d,
cutoff="iu", validate="none", show=FALSE, plot=FALSE)
s10_youden <- tabscore(event, c(age,hypertension,smoking), data=d,
cutoff="youden", validate="none", show=FALSE, plot=FALSE)
s10_sens <- tabscore(event, c(age,hypertension,smoking), data=d,
cutoff="sens", sens=.90, validate="none", show=FALSE, plot=FALSE)
s10_cost <- tabscore(event, c(age,hypertension,smoking), data=d,
cutoff="cost", cost_fn=5, cost_fp=1,
validate="none", show=FALSE, plot=FALSE)
# 11. Clinically meaningful probability threshold and multiple risk groups.
s11 <- tabscore(event, c(age,hypertension,smoking), data=d,
riskcut=c(.05,.10,.20), cutoff="risk",
validate="none", show=FALSE, plot=FALSE)
s11$tables$risk
# 12. Compare the score with an existing/reference probability.
d$reference_risk <- plogis(-4 + .04*d$age + .7*(d$hypertension == "Yes"))
s12 <- tabscore(event, c(age,hypertension,smoking), data=d,
refprob=reference_risk, refcut=.10, cutoff="refprob",
validate="none", show=FALSE, plot=FALSE)
s12$tables$reference_probability
# 13. Convert an already fitted logistic model.
m13 <- glm(event ~ age + hypertension + smoking, data=d, family=binomial())
s13 <- tabscore(m13, validate="none", show=FALSE, plot=FALSE)
# 14. External validation with a frozen scorecard.
dev <- d[1:150, ]
val <- d[151:nrow(d), ]
s14 <- tabscore(event, c(age,hypertension,smoking), data=dev,
cuts=list(age=c(40,50,60)), validation=val,
validate="none", show=FALSE, plot=FALSE)
s14$tables$external_validation
# 15. Full-pipeline bootstrap validation. B=20 is only a quick code check;
# use bootstrap=500 or more for the final report.
s15 <- tabscore(event, c(age,hypertension,smoking), data=d,
cuts=list(age=c(40,50,60)),
validate="bootstrap", bootstrap=20,
show=FALSE, plot=FALSE)
s15$tables$validation
# 16. Cox scorecard: survival is the only package required for this family.
if (requireNamespace("survival", quietly=TRUE)) {
ds <- d
true_t <- rexp(nrow(ds), rate=exp(-3 + .02*ds$age +
.6*(ds$hypertension == "Yes")))
censor_t <- rexp(nrow(ds), rate=.08)
ds$status <- as.integer(true_t <= censor_t)
ds$ftime <- pmin(true_t, censor_t)
sc <- tabscore(status, c(age,hypertension,smoking), data=ds,
family="cox", time=ftime, times=c(1,3,5), cutoff_time=5,
cuts=list(age=c(40,50,60)),
validate="none", show=FALSE, plot=FALSE)
sc$tables$risk
sc$tables$time_brier
}
# 17. Poisson count scorecard.
dp <- d
dp$count <- rpois(nrow(dp), exp(-1 + .015*dp$age + .35*(dp$smoking == "Yes")))
sp <- tabscore(count, c(age,smoking), data=dp, family="poisson",
cuts=list(age=c(40,50,60)), validate="none",
show=FALSE, plot=FALSE)
sp$tables$risk
# 18. Every available plot; Viewer includes the same figures when show=TRUE.
sv <- tabscore(event, c(age,hypertension,smoking), data=d,
cuts=list(age=c(40,50,60)), validate="none",
show=FALSE, plot=TRUE)
plot(sv, which="risk")
plot(sv, which="roc")
plot(sv, which="calibration")
plot(sv, which="decision")
plot(sv, which="distribution")
plot(sv) # all available plots in Plot history
# 19. Protective predictors: default risk-only coding versus a signed score.
dr <- d
dr$exercise <- factor(rbinom(nrow(dr), 1, .55), 0:1, c("No", "Yes"))
dr$event2 <- rbinom(nrow(dr), 1,
plogis(-2.5 + .04*dr$age - .8*(dr$exercise == "Yes")))
srisk <- tabscore(event2, c(age,exercise), data=dr,
cuts=list(age=c(40,50,60)), riskonly=TRUE,
validate="none", show=FALSE, plot=FALSE)
ssigned <- tabscore(event2, c(exercise), data=dr,
points=list(exercise=c("No"=0,"Yes"=-2)),
riskonly=FALSE, scoreref="model",
validate="none", show=FALSE, plot=FALSE)
# 20. Additional cutoff strategies.
s20_prev <- tabscore(event, c(age,hypertension,smoking), data=d,
cutoff="prevalence", validate="none",
show=FALSE, plot=FALSE)
s20_spec <- tabscore(event, c(age,hypertension,smoking), data=d,
cutoff="spec", spec=.90, validate="none",
show=FALSE, plot=FALSE)
s20_manual <- tabscore(event, c(age,hypertension,smoking), data=d,
cutoff="manual", cutoff_value=3, validate="none",
show=FALSE, plot=FALSE)
s20_none <- tabscore(event, c(age,hypertension,smoking), data=d,
cutoff="none", validate="none",
show=FALSE, plot=FALSE)
# 21. Quantile-based automatic categorization.
s21 <- tabscore(event, c(age,bmi,hypertension,smoking), data=d,
continuous="quantile", bins=4,
validate="none", show=FALSE, plot=FALSE)
# 22. An R4VN logistic() result can be converted directly as well.
rfit <- logistic(event, vars=vars(c.age, hypertension, smoking),
data=d, show=FALSE)
s22 <- tabscore(rfit, validate="none", show=FALSE, plot=FALSE)
# 23. Prediction outputs after the scorecard is frozen.
predict(s3, newp, type="score")
predict(s3, newp, type="risk")
predict(s3, newp, type="group")
predict(s3, newp, type="model")
# 24. Export all publication-ready tables without another tabscore-specific
# dependency. Word/Excel writers are only needed when those formats are chosen.
# tabexport(sv$publication_tables, export=c("html","docx","xlsx"),
# file="tabscore_report", open=FALSE)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.