predict: Prediction and postestimation diagnostics

View source: R/postestimation.R

predictR Documentation

Prediction and postestimation diagnostics

Description

predict() provides one consistent R4VN postestimation interface. With an explicit fitted model and no newvar, ordinary prediction types continue to delegate to the model's stats::predict() method. R4VN also adds common residual and influence statistics. In variable-generation mode, provide newvar and R4VN writes the selected statistic back to data or the active data frame while preserving omitted estimation rows as NA.

Usage

predict(
  object = NULL, ..., newvar = NULL, type = "auto", term = NULL,
  data = NULL, replace = FALSE, show = TRUE
)

Arguments

object

Optional fitted model or R4VN model result. If omitted when generating a variable, the most recent active model is used.

...

Additional model-specific arguments. newdata may be supplied here for ordinary explicit-model prediction and is normalized to the predictor types retained by the fitted model. Other examples include level and interval for linear-model prediction or arguments accepted by the underlying model's prediction method.

newvar

Name of a variable to create. It may be unquoted, for example newvar = stdres, or supplied as one character string. If omitted, the requested statistic is returned instead of being written to data.

type

Statistic to obtain. Common prediction aliases are "auto", "response", "fitted", "predicted", "probability"/"pr", and "link"/"xb". Residual types include "residual", "pearson", "deviance", "working", "standardized"/"stdres"/"rstandard", and "studentized"/"studres"/"rstudent". Influence statistics include "leverage"/"hat", "cooksd", "dffits", "covratio", "dfbeta", and "dfbetas". "se.fit" returns prediction standard errors. For linear models, "lower" and "upper" return confidence-limit columns. Cox models additionally support "risk", "lp", "expected", "terms", "martingale", "deviance", "score", "schoenfeld", "scaledsch", and "partial" when supported by survival.

term

Optional coefficient/term name or column number when a statistic naturally returns several columns, notably dfbeta, dfbetas, Cox score, Schoenfeld, scaled Schoenfeld, partial residuals, or term predictions. If omitted for a multi-column result, R4VN reports the available terms.

data

Data frame used for prediction and/or receiving newvar. When omitted in generation mode, the active data frame is used.

replace

Logical; allow an existing newvar to be overwritten.

show

Logical; display a short generation message. Default TRUE.

Details

Standardized residuals are computed with stats::rstandard() and studentized residuals with stats::rstudent() when those methods are available. These are different from raw residuals. For linear regression, leverage is obtained with hatvalues(), Cook's distance with cooks.distance(), DFFITS with dffits(), and COVRATIO with covratio().

Influence statistics and residuals are defined for the estimation sample. When they are written to the original active data, observations omitted from model fitting because of missing values are filled with NA.

When R4VN compact syntax declared a predictor categorical (for example i.htn) but the original data store it as numeric 0/1, prediction data are automatically reconstructed with the factor levels retained by the fitted model. Unknown new levels remain an error rather than being silently recoded.

Value

Without newvar, returns the requested prediction, residual, or diagnostic statistic. With newvar, invisibly returns the updated data frame after writing the generated variable.

Examples

# Linear regression: fitted values and regression diagnostics
d <- data.frame(
  y = c(12, 15, 17, 20, 21, 25, 28, 31, 35, 38),
  age = seq(20, 65, by = 5),
  bmi = c(19, 21, 20, 23, 25, 24, 27, 28, 30, 29)
)
usedf(d)
m1 <- regress(y, c.age, c.bmi, show = FALSE)
predict(m1, type = "response")
predict(m1, type = "standardized")
predict(m1, type = "studentized")
predict(m1, type = "leverage")
predict(m1, type = "cooksd")
predict(m1, type = "dffits")
predict(m1, type = "covratio")

# Store diagnostics in the active data frame
predict(newvar = fitted_y, type = "fitted", show = FALSE)
predict(newvar = residual_y, type = "residual", show = FALSE)
predict(newvar = stdres, type = "standardized", show = FALSE)
predict(newvar = studres, type = "studentized", show = FALSE)
predict(newvar = leverage, type = "leverage", show = FALSE)
predict(newvar = cooksd, type = "cooksd", show = FALSE)

# DFBETA/DFBETAS are coefficient-specific; select a term when generating
names(stats::coef(m1$raw$model))
predict(m1, type = "dfbetas", term = "age")
predict(m1, newvar = dfb_age, type = "dfbetas", term = "age", show = FALSE)

# Linear-model prediction standard error and confidence limits
predict(m1, type = "se.fit")
predict(m1, type = "lower", level = 0.95)
predict(m1, type = "upper", level = 0.95)

# Logistic regression: probability and residual diagnostics
g <- data.frame(
  outcome = factor(c(0,0,0,0,1,0,1,1,1,1,1,1), levels = 0:1,
                   labels = c("No", "Yes")),
  age = seq(25, 80, by = 5),
  bmi = c(20,21,22,24,23,26,25,28,29,30,31,33)
)
usedf(g)
m2 <- logistic(outcome, c.age, c.bmi, event = "Yes", show = FALSE)
predict(m2, type = "probability")
predict(m2, type = "pearson")
predict(m2, type = "deviance")
predict(m2, type = "standardized")
predict(m2, type = "leverage")
predict(m2, type = "cooksd")

R4VN documentation built on Sept. 30, 2026, 5:13 p.m.