| fast.saveload | R Documentation |
Save and load the forest component of a fitted random forest using a
directory of files. The loaded forest can be passed to
predict(). Corresponding helpers save and load lists of forests.
fast.save(o, path = NULL, testing = TRUE, units = "Mb")
fast.load(directory, path = NULL, testing = FALSE, units = "Mb")
fast.save.list(o, path = NULL, testing = FALSE, units = "Mb")
fast.load.list(directory, path = NULL, testing = FALSE, units = "Mb")
o |
For |
path |
For saving, the destination directory, including its name.
The default is |
directory |
For loading, the name of the saved forest directory
relative to |
testing |
Print component-size diagnostics using |
units |
Character string specifying the units for the diagnostic
size display, passed to |
fast.save() saves o$forest, including its retained
training data and tree information. fast.load() returns the
reconstructed forest object, which can be used directly with
predict.rfsrc.
Components outside o$forest, such as the grow object's
prediction and performance summaries, are not saved. To preserve
the complete R object instead, use saveRDS and
readRDS.
Saving requires the fst and data.table packages;
loading requires fst. Large tabular components are written
with fst::write_fst(), and the remaining forest metadata are
stored with saveRDS(..., compress = FALSE). The files use a
‘.rda’ suffix despite these different formats. Keep the entire
directory together and read it using fast.load().
Both save helpers remove the destination directory and all its contents before writing. Use a dedicated directory for the forest or forest list, rather than a directory containing other work. The destination is then created, including parent directories as needed.
Saving takes the full destination in path. Loading takes
its name in directory and its parent in path. For a
destination stored in save.path, the matching load call is
fast.load(basename(save.path), path = dirname(save.path)).
With the default save destination, use fast.load("forest").
fast.save.list() saves the elements in subdirectories
‘forest1’, ‘forest2’, and so on, inside path.
Original list names are not stored.
fast.load.list() loads every entry returned by
list.files() in the list directory and returns an unnamed
list in that order. Keep only saved forest subdirectories there.
Directory-name ordering can differ from the original numeric index
order: for example, ‘forest10’ can precede ‘forest2’.
Use fast.load() on a named subdirectory when selecting a
particular saved element.
The saved sampling-size information is
sampfrac = o$forest$sampsize(1). Loading reconstructs the
function as function(x) x * sampfrac. This preserves a
proportional sampling-size rule; it does not preserve an arbitrary
nonlinear custom function. Use saveRDS() when that original
function must be retained.
fast.saveUsed for its file-writing side effect.
Its return value is the memory-usage matrix from the final
gc(FALSE) call, not the saved forest.
fast.loadThe reconstructed forest component, retaining
its saved class, ordinarily c("rfsrc", "forest", family).
It is a forest object rather than the original grow object.
fast.save.listAn unnamed list of the return values
from the individual fast.save() calls. The forests are
saved in the destination subdirectories.
fast.load.listAn unnamed list of reconstructed forest objects in the directory-listing order described above.
rfsrc, predict.rfsrc,
saveRDS, readRDS
## ------------------------------------------------------------
## Regression: save, load, and compare predictions
## ------------------------------------------------------------
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
o <- rfsrc(mpg ~ ., data = mtcars)
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
p <- predict(o)
pp <- predict(oo)
print(summary(p$predicted - pp$predicted))
print(summary(p$predicted.oob - pp$predicted.oob))
unlink(save.path, recursive = TRUE)
}
## ------------------------------------------------------------
## Regression: a list of forests with different node sizes
## ------------------------------------------------------------
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
o1 <- rfsrc(mpg ~ ., data = mtcars, nodesize = 1)
o2 <- rfsrc(mpg ~ ., data = mtcars, nodesize = 10)
print(o1)
print(o2)
models <- list(o1, o2)
save.path <- tempfile("rfsrc-forest-list-")
invisible(fast.save.list(models, path = save.path))
oo <- fast.load.list(basename(save.path), path = dirname(save.path))
print(predict(oo[[1]]))
print(predict(oo[[2]]))
unlink(save.path, recursive = TRUE)
}
## ------------------------------------------------------------
## RFQ for imbalanced classification
## ------------------------------------------------------------
## Use matching prediction seeds when comparing class labels.
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
data(breast, package = "randomForestSRC")
dta <- na.omit(breast)
o <- imbalanced(status ~ ., data = dta, ntree = 100)
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
set.seed(19)
p <- predict(o, seed = -19)
set.seed(19)
pp <- predict(oo, seed = -19)
print(summary(p$predicted - pp$predicted))
print(summary(p$predicted.oob - pp$predicted.oob))
print(all.equal(as.character(p$class), as.character(pp$class)))
print(all.equal(as.character(p$class.oob),
as.character(pp$class.oob)))
unlink(save.path, recursive = TRUE)
}
## ------------------------------------------------------------
## Binary classification with rfq = TRUE
## ------------------------------------------------------------
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
data(breast, package = "randomForestSRC")
dta <- na.omit(breast)
o <- rfsrc(status ~ ., data = dta, rfq = TRUE, ntree = 100,
perf.type = "gmean", splitrule = "auc")
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
set.seed(19)
p <- predict(o, seed = -19)
set.seed(19)
pp <- predict(oo, seed = -19)
print(summary(p$predicted - pp$predicted))
print(summary(p$predicted.oob - pp$predicted.oob))
print(all.equal(as.character(p$class), as.character(pp$class)))
print(all.equal(as.character(p$class.oob),
as.character(pp$class.oob)))
unlink(save.path, recursive = TRUE)
}
## ------------------------------------------------------------
## Anonymous RFQ: supply the same prediction data to both forests
## ------------------------------------------------------------
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
data(breast, package = "randomForestSRC")
dta <- na.omit(breast)
o <- rfsrc.anonymous(status ~ ., data = dta, rfq = TRUE,
ntree = 100, perf.type = "gmean", splitrule = "auc")
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
set.seed(19)
p <- predict(o, newdata = dta, seed = -19)
set.seed(19)
pp <- predict(oo, newdata = dta, seed = -19)
print(summary(p$predicted - pp$predicted))
print(all.equal(as.character(p$class), as.character(pp$class)))
unlink(save.path, recursive = TRUE)
}
## ------------------------------------------------------------
## Survival
## ------------------------------------------------------------
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
data(pbc, package = "randomForestSRC")
o <- rfsrc(Surv(days, status) ~ ., data = pbc, ntree = 100)
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
set.seed(19)
p <- predict(o, seed = -19)
set.seed(19)
pp <- predict(oo, seed = -19)
print(summary(p$predicted - pp$predicted))
print(summary(p$predicted.oob - pp$predicted.oob))
unlink(save.path, recursive = TRUE)
}
## ------------------------------------------------------------
## Survival with save.memory = TRUE
## ------------------------------------------------------------
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
data(pbc, package = "randomForestSRC")
o <- rfsrc(Surv(days, status) ~ ., data = pbc,
ntree = 100, save.memory = TRUE)
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
set.seed(19)
p <- predict(o, seed = -19)
set.seed(19)
pp <- predict(oo, seed = -19)
print(summary(p$predicted - pp$predicted))
print(summary(p$predicted.oob - pp$predicted.oob))
unlink(save.path, recursive = TRUE)
}
## ------------------------------------------------------------
## Competing risks
## ------------------------------------------------------------
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
data(wihs, package = "randomForestSRC")
o <- rfsrc(Surv(time, status) ~ ., data = wihs, nsplit = 3, ntree = 100)
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
set.seed(19)
p <- predict(o, seed = -19)
set.seed(19)
pp <- predict(oo, seed = -19)
print(summary(p$predicted - pp$predicted))
print(summary(p$predicted.oob - pp$predicted.oob))
unlink(save.path, recursive = TRUE)
}
## ------------------------------------------------------------
## Multivariate regression and classification
## ------------------------------------------------------------
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
data(nutrigenomic, package = "randomForestSRC")
ydta <- data.frame(diet = nutrigenomic$diet,
genotype = nutrigenomic$genotype,
nutrigenomic$lipids)
o <- rfsrc(get.mv.formula(colnames(ydta)),
data = data.frame(ydta, nutrigenomic$genes),
ntree = 100, importance = TRUE, nsplit = 10)
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
set.seed(19)
p <- predict(o, seed = -19)
set.seed(19)
pp <- predict(oo, seed = -19)
print(summary(get.mv.predicted(p, oob = FALSE) -
get.mv.predicted(pp, oob = FALSE)))
print(summary(get.mv.predicted(p) - get.mv.predicted(pp)))
for (yn in names(p$classOutput)) {
cat("check coherence for outcome:", yn, "\n")
print(all.equal(as.character(p$classOutput[[yn]]$class),
as.character(pp$classOutput[[yn]]$class)))
print(all.equal(as.character(p$classOutput[[yn]]$class.oob),
as.character(pp$classOutput[[yn]]$class.oob)))
}
unlink(save.path, recursive = TRUE)
}
## Not run:
## ------------------------------------------------------------
## Classification: optional alzheimers data from varPro
## ------------------------------------------------------------
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
data(alzheimers, package = "varPro")
o <- rfsrc(Diagnosis ~ ., data = alzheimers)
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
set.seed(19)
p <- predict(o, seed = -19)
set.seed(19)
pp <- predict(oo, seed = -19)
print(summary(p$predicted - pp$predicted))
print(summary(p$predicted.oob - pp$predicted.oob))
print(all.equal(as.character(p$class), as.character(pp$class)))
print(all.equal(as.character(p$class.oob),
as.character(pp$class.oob)))
unlink(save.path, recursive = TRUE)
}
## ------------------------------------------------------------
## Optional memory-intensive anonymous survival test
## ------------------------------------------------------------
## This test repeats each PBC row 250 times and can require substantial memory.
if (requireNamespace("fst", quietly = TRUE) &&
requireNamespace("data.table", quietly = TRUE)) {
data(pbc, package = "randomForestSRC")
dta <- pbc[rep(seq_len(nrow(pbc)), each = 250), ]
o <- rfsrc.anonymous(Surv(days, status) ~ ., data = dta)
print(o)
save.path <- tempfile("rfsrc-forest-")
fast.save(o, path = save.path, testing = FALSE)
oo <- fast.load(basename(save.path), path = dirname(save.path))
set.seed(19)
p <- predict(o, newdata = dta, seed = -19)
set.seed(19)
pp <- predict(oo, newdata = dta, seed = -19)
print(summary(p$predicted - pp$predicted))
unlink(save.path, recursive = TRUE)
}
## End(Not run)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.