| step_nonresponse | R Documentation |
Inflates the weights of respondents to represent the nonrespondents, under the
assumption that response is ignorable given the information used. The response
propensity can be estimated by weighting classes (cells) or by a model
("propensity"), with engines ranging from logistic regression to machine
learning (regression tree, random forest, gradient boosting). Optional
K-fold cross-fitting estimates the propensity out-of-sample to avoid the
overfitting that flexible engines can introduce. The adjustment can be applied
at the person or, via cluster, the household level.
step_nonresponse(
spec,
respondent,
method = c("weighting_class", "propensity"),
by = NULL,
formula = NULL,
engine = c("logit", "tree", "forest", "boost"),
num_classes = 5L,
cluster = NULL,
crossfit = NULL,
crossfit_seed = NULL
)
spec |
a weighting_spec. |
respondent |
a 0/1 dummy column (1 = responded) or any logical condition (unquoted) TRUE for respondents. Eligible cases that are not respondents are treated as nonresponse. |
method |
"weighting_class" (cells) or "propensity" (predictive model). |
by |
character. Adjustment cells for method = "weighting_class". |
formula |
predictor formula (right-hand side only), e.g. ~ age + region, used when method = "propensity". |
engine |
engine to estimate the propensity when method = "propensity": "logit" (logistic regression, base R), "tree" (CART via package 'rpart'), "forest" (random forest via package 'ranger') or "boost" (gradient boosting via package 'xgboost'). 'rpart', 'ranger' and 'xgboost' are optional: only needed if you pick that engine. The flexible learners run with fixed default settings and their hyperparameters are not currently exposed: "tree" and "forest" use the 'rpart' and 'ranger' defaults, and "boost" uses xgboost with nrounds = 150, max_depth = 4 and eta = 0.1. |
num_classes |
integer or NULL. Controls how propensities are used: an integer forms that many propensity classes (cell adjustment within each class); NULL applies the direct factor 1/p to each unit. |
cluster |
character or NULL. If given, the adjustment is done at the cluster (e.g. household) level for whole-household nonresponse: each household counts once with its (uniform) weight; in "weighting_class" the redistribution is between responding and nonresponding households within the cells, and in "propensity" the model is fitted with one row per household (household auxiliaries), predicting the household response. The resulting factor is assigned to every member; nonresponding households go to zero. As always, only active units (weight > 0) take part, so units already dropped (unknown eligibility, ineligible) are excluded automatically. |
crossfit |
integer or NULL. If given (number of folds K >= 2), the
propensity is estimated by K-fold cross-fitting: for each fold the model is
trained on the other folds and used to predict the held-out fold, so each
unit's propensity comes from a model that did not see it. This avoids the
overfitting that flexible engines (forest, boost) can produce, which would
otherwise inflate the weights. Folds are formed by |
crossfit_seed |
integer or NULL. Seed for reproducible fold assignment
when |
The input weighting_spec with this step appended to its recipe. The
step is recorded only; it is evaluated when prep() is called.
weighting_spec(sample_survey, base_weights = pw) |>
step_nonresponse(respondent = responded, method = "weighting_class",
by = "region")
# household-level nonresponse (whole household responds or not)
weighting_spec(sample_survey, base_weights = pw) |>
step_nonresponse(respondent = responded, method = "weighting_class",
by = "region", cluster = "household_id") |>
prep()
# propensity with cross-fitting (out-of-sample, avoids overfitting)
weighting_spec(sample_survey, base_weights = pw) |>
step_nonresponse(respondent = responded, method = "propensity",
formula = ~ region + sex, engine = "logit",
num_classes = 5, crossfit = 5, crossfit_seed = 1) |>
prep()
# gradient boosting engine (requires the 'xgboost' package)
if (requireNamespace("xgboost", quietly = TRUE)) {
weighting_spec(sample_survey, base_weights = pw) |>
step_nonresponse(respondent = responded, method = "propensity",
formula = ~ region + sex + age, engine = "boost",
num_classes = 5, crossfit = 5) |>
prep()
}
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.