View source: R/plot.subsample.rfsrc.R
| plot.subsample.rfsrc | R Documentation |
Display confidence intervals for variable importance (VIMP) from a
subsample object. Select predictors, an outcome, and an
interval method using the saved replicate estimates. Prediction-error
and joint-VIMP rows can also be displayed when requested during
subsampling.
## S3 method for class 'rfsrc'
plot.subsample(x, alpha = .01, xvar.names,
standardize = TRUE, normal = TRUE, jknife = FALSE, target, m.target = NULL,
pmax = 75, main = "", sorted = TRUE, show.plots = TRUE, ...)
x |
An object returned by |
alpha |
Significance level for the displayed intervals, whose
nominal confidence level is |
xvar.names |
Names of predictor rows to display. If omitted,
all available rows are considered, subject to |
standardize |
For regression outcomes, divide VIMP by the
variance of the selected response in the full training data.
The same variance is used for all replicate estimates and any
error row. Other families are unchanged. Set |
normal |
Use normal-approximation intervals when |
jknife |
Select the delete- |
target |
For classification, an integer or class label selecting
class-specific VIMP; the default is overall VIMP. For competing
risks, an integer from 1 to |
m.target |
Name of one response in a multivariate or mixed-outcome
forest. If omitted, a default response is selected. This selects
an outcome in the saved results; it is separate from choosing
a class or event with |
pmax |
Maximum number of rows retained for display, selected by VIMP. |
main |
Main plot title. |
sorted |
Order the displayed rows by importance. |
show.plots |
Draw the plot. Set |
... |
Graphical arguments for the interval display, including
|
The default uses normal intervals and the subsampling standard
error. Use normal = TRUE, jknife = TRUE for normal intervals
with the jackknife standard error, or normal = FALSE for
nonparametric intervals. All use the estimates already stored
in x; changing alpha or the display method does not
perform another round of subsampling.
The subsampling standard error measures replicate dispersion
around the replicate mean. The jackknife calculation measures
dispersion around the full-data estimate, incorporating the
displacement between those two centers. The jackknife standard
error is not necessarily larger for a finite replicate set.
The nonparametric interval reverses empirical quantiles of the
centered subsampling roots. See subsample.rfsrc
for the calculations.
The display summarizes uncertainty in VIMP, not the distribution of the observed predictor values. Its intervals are calculated separately for each displayed statistic. A positive lower VIMP endpoint identifies a positive-importance result under that interval procedure; the display does not apply a multiple-testing correction. Prediction-error rows summarize the error itself.
xvar.names selects rows and m.target selects a response.
For a classification response or competing-risk outcome,
target additionally selects its class-specific or
event-specific statistic. These selections are made after
resampling, so one saved object can be used for several displays.
Set alpha explicitly when comparing results with
print or extract.subsample; their default is .05,
whereas the plotting default is .01.
The default display is horizontal. xlim controls its
horizontal VIMP/error axis and ylim controls its vertical
row positions. These arguments always refer to the displayed axes;
the wrapper handles the internal limit reversal used by bxp
for horizontal boxes. horizontal = FALSE gives a vertical
display, with VIMP/error on the vertical axis.
Use boxfill (or col) for the box fill,
border for its border, and whisklty and
whisklwd for interval whiskers. By default, boxes are red
when their lower endpoint is positive and blue otherwise. User
styles are retained when the intervals are redrawn above the guides.
outline = FALSE remains the default. The whiskers are
confidence limits and the boxes show the inner interval summaries,
rather than quartiles of the observed predictor values.
cex.axis, col.axis, and las control axis
labels. Set xaxt = "n" or yaxt = "n" to suppress
one axis, or axes = FALSE to suppress both. A scalar
ylab is a vertical-axis title. names supplies one
row label per selected statistic before sorting and trimming; labels
follow the same ordering as the statistics. at supplies
positions for the final displayed intervals.
show.plots = FALSE returns the plot data invisibly without
opening a graphics device. extract.subsample(x, raw = TRUE)
returns the interval matrices and replicate estimates directly.
Invisibly returns a boxplot-summary list. Its stats component
is the selected five-row confidence-interval matrix, in displayed
order, and names contains the row labels. The remaining
components are the boxplot scaffold, not additional subsampling
confidence intervals. The same list is returned with
show.plots = FALSE.
Hemant Ishwaran and Udaya B. Kogalur
Ishwaran H. and Lu M. (2019). Standard errors and confidence intervals for variable importance in random forest regression, classification, and survival. Statistics in Medicine, 38, 558-582.
Politis, D.N. and Romano, J.P. (1994). Large sample confidence regions based on subsamples under minimal assumptions. The Annals of Statistics, 22(4):2031-2050.
Shao, J. and Wu, C.J. (1989). A general theory for jackknife variance estimation. The Annals of Statistics, 17(3):1176-1197.
subsample.rfsrc, extract.subsample,
bxp
## Small settings are for illustration; increase B for final inference.
set.seed(19)
dta <- na.omit(airquality)
o <- rfsrc(Ozone ~ ., data = dta, ntree = 100,
importance = "permute", block.size = 1)
smp <- subsample(o, B = 25, verbose = FALSE)
## Three interval displays, all at the same confidence level.
plot.subsample(smp, alpha = .05, main = "Subsampling normal intervals")
plot.subsample(smp, alpha = .05, jknife = TRUE,
main = "Jackknife normal intervals")
plot.subsample(smp, alpha = .05, normal = FALSE,
main = "Nonparametric intervals")
## Restrict the display and customize the main axes and whiskers.
plot.subsample(smp, alpha = .05,
xvar.names = c("Solar.R", "Wind", "Temp"),
xlim = c(-.1, .8), las = 2, cex.axis = .75,
whisklty = 1, whisklwd = 1.5)
plot.data <- plot.subsample(smp, alpha = .05, show.plots = FALSE)
print(plot.data)
## Multivariate regression: one subsample bank, two response displays.
mv <- rfsrc(cbind(Ozone, Temp) ~ ., data = dta, ntree = 100,
importance = "permute", block.size = 1)
mv.smp <- subsample(mv, B = 25, verbose = FALSE)
plot.subsample(mv.smp, m.target = "Ozone", alpha = .05,
main = "Ozone")
plot.subsample(mv.smp, m.target = "Temp", alpha = .05,
main = "Temperature")
print(extract.subsample(mv.smp, m.target = "Temp", alpha = .05)$var.sel.Z)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.