tabmachine: Comprehensive machine-learning analysis with...

View source: R/tabmachine.R

tabmachineR Documentation

Comprehensive machine-learning analysis with publication-ready reporting

Description

tabmachine() is the comprehensive machine-learning command in R4VN. It is designed for health and biomedical research where the user needs a complete, reproducible workflow rather than only a fitted prediction model.

The function can detect the prediction task, split development data, preprocess predictors, handle missing values, encode categorical predictors, detect problematic predictors, standardize predictors when required, transform skewed numeric variables when explicitly requested, address class imbalance, select predictors, tune candidate algorithms, perform cross- validation, compare models, select a final model, determine a classification threshold using training data only, evaluate the untouched test set, calculate confidence intervals for performance measures whenever a defensible interval is implemented, assess calibration and clinical utility, calculate variable importance, and generate prediction-ready model objects.

The central R4VN principle is that automation must remain transparent. "auto" may choose an analysis action, but every action is stored in the returned object and shown in the report. Preprocessing, feature selection, class balancing, tuning, and threshold optimization are learned from training data only. The held-out test data are not used to make those decisions.

Usage

tabmachine(
  outcome,
  x = NULL,
  data = NULL,
  exclude = NULL,
  task = c("auto", "binary", "multiclass", "regression"),
  event = NULL,
  preprocess = c("auto", "none"),
  missing = c("auto", "median", "mean", "mode", "complete"),
  missing_max = 0.5,
  encode = c("auto", "dummy"),
  standardize = c("auto", "none", "z", "minmax", "robust"),
  transform = c("none", "auto", "log", "yeojohnson"),
  outlier = c("none", "detect", "winsor", "robust"),
  corr = "auto",
  feature = c("none", "auto", "polynomial", "interaction", "all"),
  degree = 2,
  reduce = c("none", "auto", "pca"),
  variance = 0.95,
  select = c("auto", "none", "filter", "lasso", "stepwise", "importance", "rfe",
    "boruta", "compare"),
  nfeatures = "auto",
  simplify = TRUE,
  simplify_tol = 0.01,
  balance = c("auto", "none", "weight", "up", "down", "smote", "adasyn", "rose",
    "compare"),
  balance_target = 0.5,
  neighbors = 5,
  split = 0.8,
  folds = 10,
  repeats = 1,
  nested = FALSE,
  seed = NULL,
  method = "auto",
  tune = TRUE,
  tune_n = 10,
  metric = "auto",
  threshold = "auto",
  target_sens = 0.9,
  target_spec = 0.9,
  ci = TRUE,
  ci_level = 0.95,
  boot = 1000,
  calibration = TRUE,
  decision = TRUE,
  decision_thresholds = seq(0.01, 0.99, 0.01),
  learning = FALSE,
  importance = TRUE,
  importance_repeats = 20,
  explain = TRUE,
  shap = FALSE,
  pdp = FALSE,
  validation = NULL,
  predict = NULL,
  id = NULL,
  show = TRUE,
  plot = TRUE,
  plot_display = "auto",
  plot_args = list(),
  strict = FALSE,
  console = FALSE,
  digit = 3,
  title = NULL,
  ai = FALSE,
  ...
)

Arguments

outcome

Outcome variable. Supply an unquoted variable name or a one-element character name. Binary, multiclass, and continuous outcomes are supported.

x

Candidate predictors. Use vars(age, sex, bmi), a character vector, a single bare variable, or . to use every variable except outcome and exclude. If omitted, all other variables are used.

data

Optional data frame. If omitted, the active R4VN data selected by usedf() or opendata(..., active = TRUE) is used.

exclude

Optional predictors to exclude, supplied as vars(...), a bare variable, or a character vector. Typical examples are identifiers, names, dates that would leak future information, or variables unavailable at the intended prediction time.

task

Prediction task: "auto" (default), "binary", "multiclass", or "regression". In auto mode, an outcome with two observed levels is binary; a factor/character outcome or a low-cardinality integer outcome is multiclass; otherwise a numeric outcome is regression.

event

Positive/event level for binary classification. If omitted, R4VN recognizes common positive encodings such as 1, TRUE, Yes, Positive, Case, or Co; otherwise the second observed level is used. The chosen event is always reported.

preprocess

Preprocessing policy. "auto" performs safe structural checks, imputation, factor-level alignment, dummy encoding, zero/near-zero variance removal, and optional correlation filtering. "none" disables automatic structural filtering but still creates a usable model matrix.

missing

Missing-value handling for predictors: "auto", "median", "mean", "mode", or "complete". "auto" uses median for numeric and mode for categorical predictors. Imputation values are estimated on each training fold and then applied to its validation fold. Outcome missingness is never imputed.

missing_max

Maximum allowed proportion missing in a candidate predictor before automatic structural filtering removes it. Default 0.50.

encode

Encoding of categorical predictors. Currently "auto" and "dummy" use treatment/dummy coding through a training-derived model.matrix blueprint. New/unseen levels in validation data are mapped to the training reference fallback and reported.

standardize

Standardization policy: "auto", "none", "z", "minmax", or "robust". In auto mode, z-standardization is applied only inside algorithms that materially benefit from scaling (penalized models, SVM, KNN, and multinomial neural optimization). Tree-based models receive the unscaled encoded matrix.

transform

Numeric transformation: "none" (default), "auto", "log", or "yeojohnson". "auto" is intentionally conservative and currently performs no transformation unless a future R4VN rule explicitly justifies one. "log" uses a training-derived shift when necessary. "yeojohnson" estimates lambda on the training data by profile likelihood.

outlier

Outlier policy: "none" (default), "detect", "winsor", or "robust". Detection uses training-data IQR rules and never deletes rows. "winsor" caps numeric values at the 1st and 99th training percentiles; "robust" caps at median +/- 5 MAD. Cut points are then reused unchanged for validation/test data.

corr

Correlation filtering for numeric candidate predictors. "auto" uses 0.95, a numeric value in (0,1) supplies a custom absolute-correlation threshold, and FALSE/"none" disables correlation filtering. Filtering is learned separately inside each training fold.

feature

Feature engineering: "none" (default), "auto", "polynomial", "interaction", or "all". "polynomial" adds powers of numeric predictors up to degree; "interaction" adds pairwise products among numeric predictors; "all" adds both. "auto" is intentionally conservative and currently behaves as "none". All engineered features are created from a fold-specific training blueprint.

degree

Highest polynomial degree for numeric feature engineering. Default 2; values 2 or 3 are supported.

reduce

Dimensionality reduction: "none" (default), "auto", or "pca". PCA is learned only on the training fold after encoding/feature engineering. "auto" uses PCA only in strongly high-dimensional settings (more encoded features than training observations and at least 50 encoded features); otherwise it remains off to preserve clinical interpretability.

variance

Target cumulative variance retained by PCA. Default 0.95.

select

Feature-selection strategy: "auto", "none", "filter", "lasso", "stepwise", "importance", "rfe", "boruta", or "compare". "auto" deliberately resolves to the native R4VN "filter" strategy so the default workflow never depends on an optional selection package. LASSO uses glmnet when explicitly requested. Boruta uses the optional Boruta package. "compare" compares available strategies by cross-validation on the development training set and does not inspect the held-out test set.

nfeatures

Number of predictors to retain for importance/RFE selection, or "auto". With "auto", R4VN chooses a compact candidate size based on the training sample size and number of available encoded predictors.

simplify

Logical; after choosing the best algorithm, search for a smaller predictor set whose development cross-validated performance is within simplify_tol of the larger selected model. This step is training- only. Default TRUE.

simplify_tol

Maximum acceptable loss in the primary metric when preferring a smaller model. For metrics where larger is better this is an absolute decrease; for RMSE/MAE it is an absolute increase. Default 0.01.

balance

Class-imbalance handling for binary classification: "auto", "none", "weight", "up", "down", "smote", "adasyn", "rose", or "compare". Balancing is applied only to the training portion of each resample. Test/validation observations are never resampled. "auto" uses class weighting when the minority class is below 20 percent and otherwise leaves the sample unchanged. "compare" compares available approaches on development cross-validation.

balance_target

Target minority proportion after sampling. Default 0.50.

neighbors

Number of nearest neighbors for native SMOTE/ADASYN. Default 5.

split

Development/test split. A single number such as 0.80 means 80% training and 20% untouched test data. FALSE uses all development data for model development (appropriate when a separate validation data set is supplied). A length-three vector such as c(.70,.15,.15) creates training, internal validation, and test partitions; the internal validation portion is combined with training for final refitting after model decisions are completed, while the test portion remains untouched.

folds

Number of cross-validation folds in the development training sample. Default 10. Classification folds are stratified when possible.

repeats

Number of repeated cross-validation repetitions. Default 1.

nested

Logical; if TRUE, hyperparameter tuning is repeated within each outer cross-validation fold. This is computationally expensive but gives a less optimistic development estimate. Regardless of this option, the final held-out test evaluation remains untouched by tuning.

seed

Optional random seed used for splitting, resampling, tuning, and bootstrap. The default NULL does not set a seed.

method

Algorithms to fit. "auto" deliberately uses a compact low-dependency R4VN core set so that a routine analysis does not require installation of an ML ecosystem: logistic regression plus a decision tree for binary outcomes, linear regression plus a decision tree for continuous outcomes, and multinomial regression plus a decision tree and R4VN-native KNN for multiclass outcomes. The native KNN fallback keeps multiclass auto usable even if recommended modelling packages are unavailable. "all" tries every supported engine that is already installed. Optional engines are never installed automatically. Methods may also be supplied explicitly: "logistic", "linear", "multinom", "lasso", "ridge", "elastic", "tree", "rf", "xgb", "svm", "knn", "naive".

tune

Hyperparameter tuning. TRUE/"grid" uses compact clinically practical grids; "random" samples candidate combinations; FALSE uses defaults. Penalized glmnet engines select lambda internally by CV.

tune_n

Maximum random-search combinations when tune = "random".

metric

Primary model-selection metric. "auto" uses ROC-AUC for most binary problems, PR-AUC when the event is uncommon, macro F1 for multiclass, and RMSE for regression. Other supported names include auc, pr_auc, accuracy, balanced_accuracy, sensitivity, specificity, f1, brier, logloss, macro_f1, weighted_f1, macro_auc, macro_pr_auc, rmse, mae, r2, and mape.

threshold

Binary classification threshold. A numeric value between 0 and 1 fixes the cutoff. "auto"/"youden" maximizes sensitivity + specificity - 1 using out-of-fold predictions from the development training set. "f1" maximizes F1. "sens" chooses the most specific threshold achieving at least target_sens; "spec" chooses the most sensitive threshold achieving at least target_spec. The test set is never used to choose the threshold.

target_sens

Target sensitivity used when threshold = "sens".

target_spec

Target specificity used when threshold = "spec".

ci

Logical; calculate 95% confidence intervals (or the level supplied by ci_level) for performance measures whenever an implemented interval is statistically meaningful. Default TRUE.

ci_level

Confidence level, default 0.95.

boot

Number of bootstrap replicates for performance measures whose interval has no preferred closed-form method. Default 1000. For final publication analyses, 2000 or more may be preferred when runtime permits.

calibration

Logical; for binary classification, calculate calibration intercept, calibration slope, Brier score, and calibration-curve data.

decision

Logical; for binary classification, calculate decision-curve net benefit for the final model over decision_thresholds.

decision_thresholds

Probability thresholds for decision-curve analysis. Default seq(.01, .99, .01).

learning

Logical; calculate training-size learning-curve summaries for the final model. Default FALSE because it can be computationally expensive.

importance

Logical; calculate permutation importance for the final model. Default TRUE. Importance is grouped back to original predictor names when dummy variables were created.

importance_repeats

Number of repeated permutations used to stabilize permutation importance. Default 20.

explain

Logical; retain explanation data and display the principal importance/calibration information in the report. Default TRUE.

shap

Logical; calculate SHAP-like contribution output when a supported engine is available. Native XGBoost predcontrib is used for XGBoost; other engines are left as NULL unless a future R4VN explainer is added.

pdp

Logical or character vector. TRUE calculates partial-dependence data for up to the five most important original numeric predictors; a character vector requests specific predictors.

validation

Optional external validation data frame. It must contain the same outcome and required predictors. It is never used for preprocessing, selection, tuning, balancing, threshold selection, or final model fitting.

predict

Optional new data frame for predictions after the final model is fitted. Predictions are returned in ⁠$predictions⁠.

id

Optional identifier variable to copy into prediction output.

show

Logical; open the publication-style HTML report in the Viewer. Default TRUE.

plot

Logical or character vector controlling figures embedded in the HTML Viewer. TRUE embeds every principal figure available for the analysis. Character values may include "roc", "pr", "calibration", "threshold", "confusion", "importance", "decision", "learning", "observed", "residual", "pdp", and "shap"; "all" requests every available figure. Any figure sent to the Plot pane through plot_display is also retained in the Viewer whenever plotting is enabled.

plot_display

Figure type(s) also drawn in the interactive R/RStudio Plot pane and Plot history. "auto" (default) draws the most useful diagnostic figures for the task. Use "all" for every available figure, "none" to keep figures only in the Viewer, or a character vector of figure names. This option never changes model fitting or performance estimates.

plot_args

Named list of base-graphics options used for Viewer and Plot pane figures. Options may be common (for example list(font_family="Arial")) or nested, for example list(all=list(font_family="Arial"), roc=list(lwd=3)).

strict

Logical. If FALSE (default), a non-essential figure that cannot be drawn is skipped with a warning while the analysis result is retained. If TRUE, such a plotting error stops the call.

console

Logical; also print a compact console summary. Default FALSE.

digit

Number of digits displayed for estimates. Default 3.

title

Optional report title.

ai

FALSE, TRUE, or a named R4VN AI endpoint. When R4VN aiask() is available, interpretation is requested only after the quantitative analysis object has been created. tabmachine() never calls AI when ai = FALSE; when enabling AI, the privacy behavior of the configured aiask() endpoint should be reviewed because fitted model objects may contain development information needed for prediction.

...

Reserved for future model-engine options.

Details

1. What tabmachine() regards as a complete ML workflow

A typical call performs the following sequence:

  1. resolve active/explicit data and variable labels;

  2. validate the outcome and candidate predictors;

  3. create an untouched test partition;

  4. inside training resamples, learn imputation/transformation/encoding rules;

  5. remove structural problems such as constant predictors;

  6. optionally create polynomial/interaction features and/or training-fold PCA;

  7. apply feature selection inside the training portion of each resample;

  8. apply class balancing only inside the training portion of each resample;

  9. tune and compare candidate algorithms;

  10. choose the final algorithm using development data only;

  11. choose a binary classification threshold from training out-of-fold predictions only;

  12. refit the selected pipeline on the full development training sample;

  13. evaluate the untouched test set and calculate confidence intervals;

  14. optionally validate on a completely external data set;

  15. assess calibration, decision-curve utility, and variable importance;

  16. store a prediction blueprint for future predict() calls.

2. Confidence intervals

ci = TRUE is the default because R4VN is intended for scientific reporting. The implementation does not attach a made-up CI to a quantity merely because a point estimate exists. Methods currently used are:

  • sensitivity, specificity, PPV, NPV, accuracy, and prevalence: Wilson binomial intervals;

  • ROC-AUC: DeLong interval through pROC when available; otherwise a stratified nonparametric bootstrap interval;

  • an optimized binary classification threshold: stratified bootstrap interval; a user-fixed threshold has no sampling CI because it is specified rather than estimated;

  • balanced accuracy, F1, MCC, kappa, PR-AUC, Brier score, log loss and other derived binary metrics: paired-observation nonparametric bootstrap;

  • calibration intercept and slope: model-based Wald intervals, with bootstrap fallback when the calibration model is unstable;

  • RMSE, MAE, R-squared and MAPE: nonparametric bootstrap over test subjects;

  • multiclass accuracy/balanced accuracy/macro-F1/weighted-F1, macro one-vs-rest AUC/PR-AUC, and log loss: nonparametric bootstrap over test subjects;

  • class-specific multiclass one-vs-rest sensitivity, specificity, PPV, NPV, accuracy and prevalence: Wilson intervals where the denominator is fixed; class-specific AUC uses DeLong through pROC when available, while PR-AUC, F1 and other derived class measures use nonparametric bootstrap;

  • decision-curve net benefit: pointwise nonparametric bootstrap when ci=TRUE.

These intervals quantify uncertainty in performance on the evaluation sample conditional on the fitted development procedure. They do not replace full external validation or transportability assessment.

3. Leakage prevention

The most important implementation rule is that no data-dependent preprocessing action is estimated on the test set. Imputation values, transformations, factor levels, scaling parameters, feature selection, balancing, tuning, and threshold optimization are fitted using training data. During CV, those operations are refitted inside each training fold before predictions are made for the corresponding validation fold.

4. Class imbalance

balance = "weight" is the preferred automatic strategy because it does not fabricate observations. up, down, native numeric-space smote, native adasyn, and optional ROSE are available for explicit experiments. Synthetic sampling occurs after fold-specific numeric encoding and is therefore applied only to the analysis portion of a resample. The original test prevalence is preserved for evaluation, PPV/NPV, calibration, Brier score and decision curves.

5. Feature selection versus feature importance

select controls which predictors are allowed into the fitted model. importance explains which predictors contribute most to a fitted final model. They are intentionally separate concepts. A variable may survive selection yet have weak final permutation importance, and correlated variables may share or exchange importance.

6. Model comparison

Development CV estimates are useful for choosing an algorithm. The untouched test estimate is the primary internal-validation result. When several models are evaluated on the same test subjects, ⁠$performance⁠ includes a CI for each available metric. ⁠$model_difference⁠ additionally stores paired bootstrap differences in the primary metric between each model and the selected model.

7. External validation

Supply validation = external_data to evaluate the finalized development pipeline without refitting it. External performance and its CIs are stored separately. If external validation is the principal evaluation, use split = FALSE to use all development observations for model development.

8. Optional packages and the low-dependency default

R4VN deliberately does not require every ML engine for every user. The default method = "auto" is intentionally based on base/recommended R engines and does not require caret, tidymodels, recipes, yardstick, or a collection of boosting/forest packages. Optional modelling engines are checked only when the user explicitly requests them or uses method = "all". glmnet supplies penalized models; ranger random forests; xgboost gradient boosting; e1071 SVM and naive Bayes; Boruta Boruta feature selection; ROSE ROSE sampling; and pROC DeLong ROC intervals. When pROC is absent, R4VN uses its native bootstrap AUC interval. The returned ⁠$engines⁠ table records what was requested, installed, and actually used.

Value

Invisibly returns an object of class c("r4vn_machine", "r4vn_tab") with major components:

overview

Data/task overview.

engines

Supported algorithms, package requirements, availability, and engines actually used.

preprocessing

Auditable preprocessing decisions.

balance

Chosen imbalance strategy and comparison when requested.

selection

Chosen feature-selection strategy and selected predictors.

tuning

Selected hyperparameters for each candidate algorithm.

cv_performance

Development cross-validation performance.

performance

Held-out test performance in long format with CI columns.

comparison

Publication-ready model comparison table.

model_difference

Paired difference in the primary metric versus the selected model, with bootstrap CI where available.

overfitting

Development-versus-evaluation comparison for the primary metric.

coefficients

For an unpenalized logistic or linear final model, model coefficients with 95% CI; logistic coefficients are exponentiated to OR.

best

Name of the selected algorithm.

final

Final fitted pipeline/model object.

threshold

Training-derived classification threshold information.

confusion

Final binary or multiclass confusion matrix counts.

class_performance

For multiclass outcomes, one-vs-rest class-specific discrimination and classification measures with 95% CI.

calibration

Calibration statistics and curve data.

decision

Decision-curve data.

importance

Grouped permutation importance.

explanation

Convenience list collecting importance, SHAP and PDP outputs when explain = TRUE.

shap

Native XGBoost contribution matrix when requested/supported.

pdp

Partial-dependence data when requested.

external_performance

External-validation performance when supplied.

external_confusion

External binary or multiclass confusion matrix when applicable.

external_class_performance

External multiclass one-vs-rest performance with CI.

external_calibration

External binary calibration statistics/curve when requested.

external_decision

External binary decision-curve data when requested.

predictions

Predictions for ⁠predict=⁠ data when supplied.

plots

Data required to replay publication plots.

plot_titles

Stable publication titles for available figure types.

tables

Named publication-ready tables used by R4VN/Studio/export.

html

Finished HTML report.

settings

Complete analysis settings.

notes

Methodological notes and any optional-engine skips.

Examples - binary classification

set.seed(123)
n <- 500
d <- data.frame(
  patient_id = seq_len(n),
  age = rnorm(n, 45, 12),
  sex = factor(sample(c("Female", "Male"), n, TRUE)),
  bmi = rnorm(n, 23, 3.5),
  smoke = factor(sample(c("No", "Yes"), n, TRUE, c(.75, .25)))
)
lp <- -6 + .055*d$age + .10*d$bmi + .65*(d$smoke == "Yes")
d$hypertension <- factor(rbinom(n, 1, plogis(lp)), 0:1, c("No", "Yes"))

m <- tabmachine(
  hypertension,
  x = vars(age, sex, bmi, smoke),
  data = d,
  event = "Yes",
  method = "logistic",
  tune = FALSE,
  boot = 200,
  show = FALSE,
  plot = FALSE
)
m$comparison
m$performance

Examples - automatic ML comparison

\donttest{
# The automatic comparison uses the low-dependency R4VN core set.
# Optional engines join only when explicitly requested or method = "all".
m <- tabmachine(
  hypertension,
  x = vars(age, sex, bmi, smoke),
  data = d,
  event = "Yes",
  method = "auto",
  select = "auto",
  balance = "auto",
  ci = TRUE,
  boot = 2000
)
plot(m, "roc")
plot(m, "calibration")
plot(m, "importance")
}

Examples - class imbalance and feature selection

\donttest{
m2 <- tabmachine(
  hypertension,
  x = .,
  data = d,
  exclude = vars(patient_id),
  event = "Yes",
  balance = "compare",
  select = "compare",
  metric = "pr_auc",
  nested = TRUE
)
m2$balance
m2$selection
}

Examples - continuous outcome

set.seed(321)
r <- data.frame(
  age = rnorm(400, 50, 14),
  bmi = rnorm(400, 24, 4),
  sex = factor(sample(c("Female", "Male"), 400, TRUE))
)
r$sbp <- 75 + .65*r$age + 1.15*r$bmi + 4*(r$sex == "Male") + rnorm(400, 0, 9)

mr <- tabmachine(
  sbp,
  x = vars(age, bmi, sex),
  data = r,
  method = "linear",
  tune = FALSE,
  boot = 200,
  show = FALSE,
  plot = FALSE
)
mr$performance

Examples - active data and new predictions

\donttest{
usedf(d)
fit <- tabmachine(
  hypertension,
  x = vars(age, sex, bmi, smoke),
  event = "Yes",
  method = "logistic"
)

newpatients <- data.frame(
  age = c(35, 68),
  sex = factor(c("Female", "Male"), levels = levels(d$sex)),
  bmi = c(21, 31),
  smoke = factor(c("No", "Yes"), levels = levels(d$smoke))
)
predict(fit, newpatients)
}

Examples - external validation

\donttest{
dev <- d[1:350, ]
ext <- d[351:500, ]
me <- tabmachine(
  hypertension,
  x = vars(age, sex, bmi, smoke),
  data = dev,
  validation = ext,
  split = FALSE,
  event = "Yes",
  method = "auto"
)
me$external_performance
}

Examples - Viewer and Plot pane together

\donttest{
# Every requested figure remains in the Viewer. The selected figures are
# also added to Plot history so Previous/Next can be used in RStudio.
mv <- tabmachine(
  hypertension, vars(age, sex, bmi, smoke), data = d, event = "Yes",
  method = "logistic", tune = FALSE, boot = 200,
  plot = TRUE,
  plot_display = c("roc", "calibration", "confusion", "importance")
)
names(mv$plots)
mv$settings$viewer_plots
mv$settings$display_plots

# Keep all figures in Viewer but draw none in the Plot pane.
mv2 <- tabmachine(
  hypertension, vars(age, sex, bmi, smoke), data = d, event = "Yes",
  method = "logistic", tune = FALSE, boot = 200,
  plot = TRUE, plot_display = "none"
)
}

Examples - classification thresholds

\donttest{
myouden <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method="logistic", tune=FALSE, threshold="youden",
  boot=200, show=FALSE, plot=FALSE)
mf1 <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method="logistic", tune=FALSE, threshold="f1",
  boot=200, show=FALSE, plot=FALSE)
msens <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method="logistic", tune=FALSE, threshold="sens",
  target_sens=.90, boot=200, show=FALSE, plot=FALSE)
mspec <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method="logistic", tune=FALSE, threshold="spec",
  target_spec=.90, boot=200, show=FALSE, plot=FALSE)
mfixed <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method="logistic", tune=FALSE, threshold=.20,
  boot=200, show=FALSE, plot=FALSE)
myouden$threshold
plot(myouden, "threshold")
}

Examples - missing data, correlation, outliers, and transformations

\donttest{
mp <- tabmachine(
  hypertension, vars(age, sex, bmi, smoke), data=d, event="Yes",
  method="logistic", tune=FALSE,
  missing="auto", corr=.90, outlier="winsor", transform="none",
  boot=200, show=FALSE, plot=FALSE
)
mp$preprocessing
}

Examples - imbalance without extra packages

\donttest{
# weight, up, down, SMOTE, and ADASYN are implemented inside R4VN.
mw <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method="logistic", tune=FALSE, balance="weight",
  metric="pr_auc", boot=200, show=FALSE, plot=FALSE)
msmote <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method="logistic", tune=FALSE, balance="smote",
  metric="pr_auc", boot=200, show=FALSE, plot=FALSE)
mw$balance
msmote$balance
}

Examples - feature engineering and PCA

\donttest{
mfeat <- tabmachine(hypertension, vars(age, bmi), data=d, event="Yes",
  method="logistic", tune=FALSE, feature="all", degree=2,
  boot=200, show=FALSE, plot=FALSE)
mfeat$selection

mpdp <- tabmachine(hypertension, vars(age, bmi, smoke), data=d, event="Yes",
  method="logistic", tune=FALSE, pdp=c("age","bmi"), boot=200,
  plot=TRUE, plot_display="pdp")
mpdp$pdp
plot(mpdp, "pdp")

set.seed(11)
hd <- as.data.frame(matrix(rnorm(180*60),180,60))
names(hd) <- paste0("x",1:60)
hd$y <- factor(rbinom(180,1,plogis(hd$x1-.7*hd$x2+.5*hd$x3)),0:1,c("No","Yes"))
mpca <- tabmachine(y, x=., data=hd, event="Yes", method="logistic",
  tune=FALSE, reduce="pca", variance=.90, select="filter",
  simplify=FALSE, boot=100, show=FALSE, plot=FALSE)
mpca$selection
}

Examples - feature selection and parsimony

\donttest{
mfilter <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method="logistic", tune=FALSE, select="filter",
  simplify=TRUE, boot=200, show=FALSE, plot=FALSE)
mfilter$selection
mfilter$simplify

# Penalized selection is optional and used only when glmnet is installed.
if (requireNamespace("glmnet", quietly=TRUE)) {
  mlasso <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
    event="Yes", method="lasso", select="lasso", boot=200,
    show=FALSE, plot=FALSE)
  mlasso$selection
}
}

Examples - repeated and nested validation

\donttest{
mrep <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method=c("logistic","tree"), tune=TRUE,
  folds=5, repeats=3, boot=200, show=FALSE, plot=FALSE)
mrep$cv_performance

mnested <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method=c("logistic","tree"), tune=TRUE,
  folds=5, nested=TRUE, boot=200, show=FALSE, plot=FALSE)
mnested$cv_performance
}

Examples - three-way split

\donttest{
msplit <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", split=c(.70,.15,.15), method="auto", boot=200,
  show=FALSE, plot=FALSE)
msplit$overview
}

Examples - calibration, clinical utility, and learning curve

\donttest{
mclin <- tabmachine(hypertension, vars(age, sex, bmi, smoke), data=d,
  event="Yes", method="logistic", tune=FALSE,
  calibration=TRUE, decision=TRUE, learning=TRUE, boot=200,
  plot=TRUE, plot_display="all")
mclin$calibration$statistics
head(mclin$decision)
mclin$learning
}

Examples - regression diagnostics

\donttest{
mr2 <- tabmachine(sbp, vars(age,bmi,sex), data=r, method="auto",
  boot=200, plot=TRUE, plot_display=c("observed","residual","importance"))
mr2$comparison
plot(mr2,"observed")
plot(mr2,"residual")
}

Examples - multiclass classification

\donttest{
ir <- iris
mm <- tabmachine(Species, vars(Sepal.Length,Sepal.Width,Petal.Length,Petal.Width),
  data=ir, method="auto", folds=5, boot=200,
  plot=TRUE, plot_display=c("confusion","importance"))
mm$performance
mm$class_performance
mm$confusion
plot(mm,"confusion")
}

Examples - predictions for new patients

\donttest{
newpatients <- data.frame(
  patient_id=c("P001","P002"), age=c(35,68),
  sex=factor(c("Female","Male"),levels=levels(d$sex)),
  bmi=c(21,31), smoke=factor(c("No","Yes"),levels=levels(d$smoke))
)
mpred <- tabmachine(hypertension, vars(age,sex,bmi,smoke), data=d,
  event="Yes", method="logistic", tune=FALSE, boot=200,
  predict=newpatients, id=patient_id, show=FALSE, plot=FALSE)
mpred$predictions
predict(mpred,newpatients,type="prob")
predict(mpred,newpatients,type="class")
}

Examples - optional advanced engines

\donttest{
# None of these packages is needed for the ordinary R4VN auto workflow.
if (requireNamespace("ranger",quietly=TRUE)) {
  mrf <- tabmachine(hypertension, vars(age,sex,bmi,smoke), data=d,
    event="Yes", method="rf", boot=200, show=FALSE, plot=FALSE)
}
if (requireNamespace("xgboost",quietly=TRUE)) {
  mxgb <- tabmachine(hypertension, vars(age,sex,bmi,smoke), data=d,
    event="Yes", method="xgb", shap=TRUE, boot=200, show=FALSE, plot=FALSE)
  mxgb$shap
  plot(mxgb, "shap")
}
if (requireNamespace("e1071",quietly=TRUE)) {
  msvm <- tabmachine(hypertension, vars(age,sex,bmi,smoke), data=d,
    event="Yes", method="svm", boot=200, show=FALSE, plot=FALSE)
}
if (requireNamespace("Boruta",quietly=TRUE)) {
  mb <- tabmachine(hypertension, vars(age,sex,bmi,smoke), data=d,
    event="Yes", method="logistic", select="boruta", tune=FALSE,
    boot=200, show=FALSE, plot=FALSE)
}
if (requireNamespace("ROSE",quietly=TRUE)) {
  mrose <- tabmachine(hypertension, vars(age,sex,bmi,smoke), data=d,
    event="Yes", method="logistic", balance="rose", tune=FALSE,
    boot=200, show=FALSE, plot=FALSE)
}
}

Examples

set.seed(99)
n <- 110
dd <- data.frame(
  x1 = rnorm(n),
  x2 = rnorm(n),
  group = factor(sample(c("A", "B"), n, TRUE))
)
pp <- plogis(-0.4 + 0.9 * dd$x1 - 0.5 * dd$x2)
dd$y <- factor(rbinom(n, 1, pp), 0:1, c("No", "Yes"))

z <- tabmachine(
  y,
  x = vars(x1, x2, group),
  data = dd,
  event = "Yes",
  method = "logistic",
  tune = FALSE,
  folds = 2,
  ci = FALSE,
  boot = 50,
  calibration = FALSE,
  decision = FALSE,
  importance = FALSE,
  explain = FALSE,
  show = FALSE,
  plot = FALSE
)
z$comparison

R4VN documentation built on Sept. 30, 2026, 5:13 p.m.