covselda: CovSel-discriminant analysis

covselrdaR Documentation

CovSel-discriminant analysis

Description

Variable selection for high-dimensionnal data with the COVSEL method (Roger et al. 2011), followed by a linear regression model.

The training variable y (univariate class membership) is firstly transformed to a dummy table containing nclas columns, where nclas is the number of classes present in y. Each column is a dummy variable (0/1). Then, a variable selection, based on the COVSEL method, is implemented on the X-data and the dummy table, returning a set of X-variables that are used as dependent variables in a DA model.

- covselrda: A linear regression model predicts the Y-dummy table from the selected X-variables. For a given observation, the final prediction is the class corresponding to the dummy variable for which the prediction is the highest.

- covsellda and covselqda: Probabilistic LDA and QDA are run over the selected X-variables, respectively.

Auxiliary functions

predict Calculates the predictions for any new set of variables contained in the selection.

Usage


covselrda(X, y, nvar = NULL, Xscaling = c("none", "pareto", "sd")[1], 
Yscaling = c("none", "pareto", "sd")[1], weights = NULL)

covsellda(X, y, nvar = NULL, prior = c("unif", "prop"), 
Xscaling = c("none", "pareto", "sd")[1], 
Yscaling = c("none", "pareto", "sd")[1], weights = NULL)
  
covselqda(X, y, nvar = NULL, prior = c("unif", "prop"), 
Xscaling = c("none", "pareto", "sd")[1], 
Yscaling = c("none", "pareto", "sd")[1], weights = NULL)

## S3 method for class 'Covselrda'
predict(object, X, ..., nvar = NULL)

## S3 method for class 'Covselprobda'
predict(object, X, ..., nvar = NULL)

Arguments

X

X-data (n, p).

y

Training class membership (n). Note: If y is a factor, it is replaced by a character vector.

nvar

Number of variables to select in X. Can be a vector for the auxiliary functions

prior

The prior probabilities of the classes. Possible values are "unif" (default; probabilities are set equal for all the classes) or "prop" (probabilities are set equal to the observed proportions of the classes in y).

Xscaling

X variable scaling among "none" (mean-centering only), "pareto" (mean-centering and pareto scaling), "sd" (mean-centering and unit variance scaling). If "pareto" or "sd", uncorrected standard deviation is used.

Yscaling

Y variable scaling among "none" (mean-centering only), "pareto" (mean-centering and pareto scaling), "sd" (mean-centering and unit variance scaling). If "pareto" or "sd", uncorrected standard deviation is used.

weights

Weights (n, 1) to apply to the training observations. Internally, weights are "normalized" to sum to 1. Default to NULL (weights are set to 1 / n).

object

For the auxiliary functions: A fitted model, output of a call to the main functions.

...

For the auxiliary functions: Optional arguments. Not used.

Value

For covselrda, covselda, covselqda:

sel

A dataframe where variable sel shows the column indexes of the variables selected in X.

fm

List of linear regression or discriminant models, involving 1 to nvar selected explicative variables.

lev

classes

ni

number of observations in each class

weights

The weights used for the row observations.

For predict.Covselrda, predict.Covselprobda:

pred

predicted class for each observation

posterior

calculated probability of belonging to a class for each observation

References

Roger, J.M., Palagos, B., Bertrand, D., Fernandez-Ahumada, E., 2011. CovSel: Variable selection for highly multivariate and multi-response calibration: Application to IR spectroscopy. Chem. Lab. Int. Syst. 106, 216-223.

Examples


n <- 6 ; p <- 4
X <- matrix(rnorm(n * p), ncol = p)
y <- c("A","A","B","B","C","C")

sel <- covselrda(X, y, nvar = 3)

predict(sel, X, nvar = c(2,3))


rchemo documentation built on June 30, 2026, 5:10 p.m.