| gg_partial | R Documentation |
A partial dependence curve answers a what-if question about a forest: hold every other predictor at its observed value, sweep one of them across its range, and watch how the ensemble prediction moves. Marginalized over the joint distribution of the other variables, the resulting curve isolates the average effect of the swept predictor alone.
gg_partial(part_dta, nvars = NULL, cat_limit = 10, model = NULL)
part_dta |
partial plot data from
|
nvars |
how many of the partial plot variables to calculate |
cat_limit |
Categorical features are built when there are fewer than
|
model |
a label name applied to all features. Useful when combining multiple partial plot objects in figures. |
gg_partial handles the bookkeeping step after you've already called
randomForestSRC::plot.variable(partial = TRUE): it takes the list
that function returns and separates the variables into two tidy data frames
– one for continuous predictors (plotted as lines) and one for categorical
predictors (plotted as bar charts). The split is controlled by
cat_limit: variables with more unique x-values than this threshold
are treated as continuous; all others are categorical.
If you'd rather skip the plot.variable step and pass the fitted
forest directly, see gg_partial_rfsrc, which calls
partial.rfsrc for you.
A named list with two elements:
data.frame with columns x, yhat,
name (and optionally model) for continuous variables
data.frame with the same columns but with x
as a factor, for low-cardinality / categorical variables
Partial-dependence extraction is randomForestSRC-only;
there is no randomForest method (the randomForest package
provides no comparable partial-dependence interface).
For survival forests, randomForestSRC::plot.variable defaults
to surv.type = "mort", so yhat is mortality – the
expected number of events – and not a survival probability. It is
therefore not on [0, 1] and is not directly comparable with the
survival probabilities returned by gg_variable. For a
comparable quantity, ask for it explicitly:
randomForestSRC::plot.variable(rf, partial = TRUE,
surv.type = "surv"). The label describing the plotted quantity is
recorded on the returned object as attr(x, "ylabel") and is used
as the y-axis title by plot.gg_partial.
Note that gg_partial_rfsrc defaults to
partial.type = "surv" and so reports survival probabilities. The
two entry points therefore report different quantities by default.
gg_partial_rfsrc gg_partialpro
## Build a small regression forest on the airquality dataset
set.seed(42)
airq <- na.omit(airquality)
rf <- randomForestSRC::rfsrc(Ozone ~ ., data = airq, ntree = 50)
## Compute partial dependence via plot.variable (show.plots = FALSE to
## suppress the base-graphics output, we only want the data)
pv <- randomForestSRC::plot.variable(rf, partial = TRUE,
show.plots = FALSE)
## Split into continuous and categorical data frames
result <- gg_partial(pv)
head(result$continuous)
## Label this model for later comparison with a second forest
result_labeled <- gg_partial(pv, model = "airq_model")
unique(result_labeled$continuous$model)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.