| lava | R Documentation |
Adjust a lava regularized linear model, that is a lava transformation
of the data followed by a
(possibly weighted) \ell_1-norm. The solution path is
computed at a grid of values for the \ell_1-penalty. See
details for the criterion optimized.
lava(
x,
y,
lambda1 = NULL,
lambda2 = 1,
weights = rep(1, nrow(x)),
penscale = rep(1, ncol(x)),
struct = Matrix::Diagonal(ncol(x), 1),
intercept = TRUE,
normalize = TRUE,
refit = FALSE,
nlambda1 = ifelse(is.null(lambda1), 100, length(lambda1)),
minratio = 0.01,
maxfeat = min(nrow(x), ncol(x)),
beta0 = numeric(ncol(x)),
control = list()
)
x |
matrix of features, possibly sparsely encoded
(experimental). Do NOT include intercept. When normalized is
|
y |
response vector. |
lambda1 |
sequence of decreasing |
lambda2 |
real scalar; tunes the |
weights |
vector with real positive values that weight the observations (like in weighted least square). Default sets all weights to 1. |
penscale |
vector with real positive values that weight the penalty of each feature. Default sets all weights to 1. |
struct |
matrix structuring the coefficients, possibly
sparsely encoded. Must be at least positive semidefinite (this is
checked internally). If |
intercept |
logical; indicates if an intercept should be
included in the model. Default is |
normalize |
logical; indicates if variables should be
normalized to have unit L2 norm before fitting. Default is
|
refit |
logical: indicates if the non null coefficients should be refit to avoid excessive bias. Default is FALSE. Can be changed later (both raw and refit coefficients are stored). |
nlambda1 |
integer that indicates the number of values to put
in the |
minratio |
minimal value of |
maxfeat |
integer; limits the number of features ever to
enter the model; i.e., non-zero coefficients for the Elastic-net:
the algorithm stops if this number is exceeded and |
beta0 |
a starting point for the vector of parameter. By default, will initialized zero. May save time in some situation. |
control |
list of argument controlling low level options of the algorithm –use with care and at your own risk– :
|
The optimized criterion is the following: βhat λ1 = argminθ = β+δ 1/2n RSS(β + δ) + λ1 | β |1 + λ/2 2 δT S δ, .
an object with class QuadrupenFit.
an object with class LavaFit, inheriting from QuadrupenFit.
Chernozhukov, Victor, Christian Hansen, and Yuan Liao. "A lava attack on the recovery of sums of dense and sparse signals." The Annals of Statistics (2017): 39-76. doi:10.1214/16-AOS1434
## Simulating multivariate Gaussian with blockwise correlation
## and piecewise constant vector of parameters
beta <- rep(c(0,1,0,-1,0), c(25,10,25,10,25))
delta <- runif(sum(c(25,10,25,10,25)),-.1,.1)
cor <- 0.75
Soo <- toeplitz(cor^(0:(25-1))) ## Toeplitz correlation for irrelevant variables
Sww <- matrix(cor,10,10) ## bloc correlation between active variables
Sigma <- Matrix::bdiag(Soo,Sww,Soo,Sww,Soo)
diag(Sigma) <- 1
n <- 50
x <- as.matrix(matrix(rnorm(95*n),n,95) %*% chol(Sigma))
y <- 10 + x %*% beta + rnorm(n,0,10)
labels <- rep("irrelevant", length(beta))
labels[beta != 0] <- "relevant"
## The solution path of the LAVA
out <- lava(x,y)
out$plot_path(component = "sparse", labels=labels)
out$plot_path(component = "dense", labels=labels)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.