pls.single.cv: Single cross-validation for PLS component optimization

View source: R/main.R

pls.single.cvR Documentation

Single cross-validation for PLS component optimization

Description

Performs grouped k-fold or leave-one-out cross-validation over candidate component counts and, when vector-valued predictive arguments are supplied, over a compact hyperparameter grid. Selection uses a task-appropriate metric returned by evaluate() or the fitted-response R2Y path.

Usage

pls.single.cv(
  Xdata,
  Ydata,
  ncomp = 2,
  constrain = NULL,
  scaling = c("centering", "autoscaling", "none"),
  method = c("simpls", "plssvd", "opls", "kernelpls"),
  backend = NULL,
  n.cores = NULL,
  seed = 1L,
  kfold = 10,
  north = 1L,
  kernel = c("linear", "rbf", "poly"),
  gamma = NULL,
  degree = 3L,
  coef0 = 1,
  classifier = c("lda", "argmax"),
  fit = TRUE,
  bycol = FALSE,
  selection = "auto",
  return_splits = FALSE,
  ...
)

Arguments

Xdata

Predictor matrix.

Ydata

Response. Use a numeric vector/matrix for regression or factor/character class labels for classification.

ncomp

Positive integer component count or vector of counts. Repeated values are removed. When a direct PLS-LDA fit reaches its numerical rank, all requested positions are retained and positions beyond that rank repeat the last estimable prediction and discriminant scores. The fitted object reports both requested and effective component counts.

constrain

Optional grouping vector for grouped cross-validation. It must have one value per sample. Samples with the same value are assigned to the same fold, so all rows from the same patient, subject, batch, or technical replicate stay together in training or test data. When NULL, each sample is treated as its own group.

scaling

One of centering, autoscaling, or none.

method

One or more of simpls, plssvd, opls, or kernelpls. Multiple values are treated as a tuning grid. simpls denotes the fastPLS SIMPLS-family estimator; an eligible bounded-block route is approximate rather than classical de Jong SIMPLS.

backend

Implementation backend: cpu, cuda, or metal. Multiple values are treated as a tuning grid. Metal requires float32 input and uses fixed CPU/Metal operation splitting with CPU fold orchestration and prediction. When omitted, options(backend = ...) defines the session default. Every requested backend must be available; unavailable accelerator entries raise an error instead of using CPU.

n.cores

Number of CPU cores requested for supported BLAS/OpenMP host operations. An explicit value takes precedence over options(n.cores = ...). CUDA and Metal device parallelism is managed by their native runtimes; n.cores applies only to their host-side stages.

seed

Random seed used for fold assignment and randomized SVD steps.

kfold

Number of folds, or "loocv" for leave-one-out cross-validation. When constrain is supplied, LOOCV means leave-one-constraint-group-out: samples sharing the same constraint value are always held out together and are never split across training and test. A numeric value at least as large as the number of groups has the same leave-one-group-out interpretation.

north

Number of orthogonal components removed by OPLS. The predictive count ncomp must fit within the predictor rank remaining after filtering. Direct PLS-LDA fits retain larger requested positions by repeating the last estimable result; other OPLS routes reject an unavailable predictive direction. Orthogonal components are not counted as predictive components.

kernel

Kernel type for kernel PLS: linear, rbf, or poly.

gamma

Kernel scale. Defaults internally to 1 / ncol(Xdata). For method = "kernelpls", multiple values are treated as a tuning grid.

degree

Polynomial kernel degree.

coef0

Polynomial kernel offset.

classifier

Classification rule for factor responses. The default "lda" fits linear discriminant analysis in the latent score space; "argmax" selects the largest PLS-DA response score. Multiple values are treated as a tuning grid.

fit

Fit one additional model on the full dataset and return its fitted values (Yfit) and training R2Y path. The default is TRUE for backward compatibility. Set to FALSE to skip this extra full-data fit; held-out cross-validated Q2Y and RMSD are still calculated. Selecting R2Y overrides fit = FALSE because fitted responses are required for that criterion.

bycol

For matrix-valued regression responses, calculate response-wise metrics in the returned metrics list. The default FALSE returns only aggregate metrics.

selection

Metric used to select settings. "auto" uses accuracy for classification and RMSD for regression. Classification also supports "balanced_accuracy", "AUROC", "lift_accuracy", "macro_precision", "macro_recall", "macro_f1", "kappa", "R2Y", and "Q2Y". Regression also supports "R2Y", "Q2Y", "RMSD", "MAE", "MAPE_percent", "RPD", "Pearson_r", and "Spearman_r". R2Y uses a model fitted to the complete supplied dataset and therefore forces fit = TRUE; as a training criterion, it is usually less suitable for complexity selection than a held-out criterion. Q2Y uses out-of-fold predictions and each fold's training-response mean. The AUROC is available for binary responses and pools continuous held-out class scores across all folds, using the second factor level as the positive class. The remaining metrics use the aggregate definitions in evaluate() on the out-of-fold predictions. RMSD, MAE, and MAPE_percent are minimized; all other criteria are maximized. Classification R2Y and Q2Y operate on dummy-coded responses, not decoded labels. Incompatible task/metric combinations raise an error before fitting. The former names "r2" and "q2" are rejected as ambiguous.

return_splits

Return a sample-by-fold character matrix named split_index when TRUE. Rows use the original sample positions and each column represents one fold, with values "training" or "test". The default FALSE avoids allocating this additional matrix.

...

Optional SVD tuning controls forwarded to the selected backend. Use the same compact names documented in fastsvd(), such as oversample and power. Vector values are included in the tuning grid.

Details

For LDA classification, each training fold uses only the classes represented in that fold and maps predictions back to the original factor levels. A class absent from a fold's training data cannot be predicted in that fold; its held-out observations remain in the reported metrics. When otherwise identical "argmax" and "lda" configurations are tuned together, the PLS fit and fold projection are calculated once; each classifier still receives its own predictions, metric path, and selected component count. For sufficiently tall SIMPLS classification problems, full-data predictor and class moments are calculated once and each fold's training moments are obtained by subtracting its held-out contribution. The fold-specific centering, scaling, PLS fit, LDA fit, and predictions remain independent. In either task, a fold that estimates fewer latent components than requested uses its available component prefix, and higher requested prefixes repeat the last estimable prediction and metric. In regression, a zero-component fold predicts its training-response mean. In LDA classification, a zero-component fold uses the empirical training-class priors and finite log-prior discriminant scores; a single-class training fold predicts its represented class. The effective_ncomp matrix records the component count used for every fold and requested prefix, while status identifies regular, zero-direction, and single-class folds. The returned tuning path always retains every requested component count.

Value

A list describing the cross-validation run and selected model. metrics$cross_validated contains complete evaluate() results for each requested component count and metrics$fitted contains the corresponding full-data fit results when fit = TRUE. metrics$definitions records the exact R2Y and Q2Y denominator conventions. Other fields are:

  • best_ncomp: number of components selected by the chosen metric.

  • best_index: position of best_ncomp in the tested component grid.

  • selection_metric: metric used for optimization. With "auto", classification uses accuracy and regression uses the default prediction error rule.

  • best_metric_name and best_metric_value: name and value of the metric at the selected component count.

  • Q2Y: held-out cross-validated Q2; every held-out fold is centered on its corresponding fold-training response mean. For factor responses, this is dummy-response PLS-DA Q2 using fold-training class proportions and is not classification accuracy. Values are named by component count.

  • accuracy: held-out decoded-label accuracy for factor responses, named by component count.

  • balanced_accuracy: held-out mean class recall for factor responses, named by component count.

  • RMSD: held-out root mean squared deviation for regression. It is NA for classification. Values are named by component count.

  • Yfit: fitted values from the full-data model when fit = TRUE.

  • R2Y: training-set explained-variance path from a model fitted on the full dataset when fit = TRUE; otherwise NA. For factor responses, this is calculated on the dummy-coded PLS-DA response scores, not on the decoded class labels.

  • fold: fold assignment used for each sample.

  • pred: decoded cross-validated predictions when predictions are stored.

  • Ypred: raw prediction array when score predictions are stored.

  • lda_scores: held-out LDA discriminant-score array when LDA scores are stored. For LDA, dummy-response PLS scores are reduced into Q2Y while each fold is processed rather than retained as a second full score array.

  • effective_ncomp: integer matrix with one row per fold and one column per requested component count. Each entry is the estimable prefix used in that fold. Zero identifies the class-prior fallback.

  • status: fold status vector. Status 1 is a regular fit, 4 is a single-class training-fold fallback, and 5 is a zero-direction class-prior fallback.

  • degenerate_inner_folds and constant_classifier_fallback: identifiers of single-class training folds and whether the constant-class prediction rule was used.

  • minimum_positive_training_count and minimum_negative_training_count: smallest binary-class counts among fold-training partitions. The second factor level is positive.

  • component_selection_informative: whether at least one fold could estimate discrimination and the selected metric path contained a finite value.

  • metrics: complete evaluate() outputs. cross_validated contains one result per requested component count from held-out predictions and fitted contains full-data fit results when fit = TRUE. For multivariate regression, response-wise metrics are included only when bycol = TRUE.

  • selection_metrics: compact per-component metric table used for component selection by the CV backend.

  • best_parameters: compact list containing only ncomp plus the arguments that were actually optimized, for example classifier when classifier = c("lda", "argmax").

  • tuning_config: relevant selected configuration used for the run. Irrelevant classifier- or method-specific defaults are omitted; for example, controls belonging to an unselected classifier are omitted.

  • tuning_summary and tuning_metrics: tables for all tested configurations when more than one predictive configuration is supplied.

  • split_index: sample-by-fold training/test membership matrix, returned only when return_splits = TRUE.

Examples

idx <- c(seq_len(12), 51:62, 101:112)
X <- as.matrix(iris[idx, seq_len(4)])
y <- factor(iris[idx, 5])
opt <- pls.single.cv(X, y,
    ncomp = seq_len(2), kfold = 3, method = "simpls",
    backend = "cpu", seed = 1
)
opt$best_ncomp
opt_kernel <- pls.single.cv(X, y,
    ncomp = seq_len(2), kfold = 3,
    method = "kernelpls", backend = "cpu",
    kernel = c("linear", "rbf"),
    gamma = c(0.1, 1), seed = 1
)
opt_kernel$best_parameters

fastPLS documentation built on Sept. 29, 2026, 1:06 a.m.