pls.double.cv: Nested cross-validation for PLS

View source: R/main.R

pls.double.cvR Documentation

Nested cross-validation for PLS

Description

Performs nested grouped cross-validation with an outer loop for unbiased performance estimation and an inner loop for component and hyperparameter selection. Constraint groups are respected in both loops so related samples remain in the same fold.

Usage

pls.double.cv(
  Xdata,
  Ydata,
  ncomp = 2,
  constrain = seq_len(nrow(Xdata)),
  scaling = c("centering", "autoscaling", "none"),
  method = c("simpls", "plssvd", "opls", "kernelpls"),
  backend = NULL,
  n.cores = NULL,
  seed = 1L,
  perm.test = FALSE,
  times = 100,
  runn = 1,
  kfold_inner = 10,
  kfold_outer = 10,
  north = 1L,
  kernel = c("linear", "rbf", "poly"),
  gamma = NULL,
  degree = 3L,
  coef0 = 1,
  classifier = c("lda", "argmax"),
  bycol = FALSE,
  selection = "auto",
  return_splits = FALSE,
  ...
)

Arguments

Xdata

Predictor matrix.

Ydata

Response. Use a numeric vector/matrix for regression or factor/character class labels for classification.

ncomp

Positive integer component count or vector of counts. Repeated values are removed. When a direct PLS-LDA fit reaches its numerical rank, all requested positions are retained and positions beyond that rank repeat the last estimable prediction and discriminant scores. The fitted object reports both requested and effective component counts.

constrain

Grouping vector for grouped cross-validation. It must have one value per sample. Samples with the same value are assigned to the same fold, so all rows from the same patient, subject, batch, or technical replicate stay together in training or test data. The default seq_len(nrow(Xdata)) treats every sample as an independent group.

scaling

One of centering, autoscaling, or none.

method

One or more of simpls, plssvd, opls, or kernelpls. Multiple values are tuned in the inner loop. simpls denotes the fastPLS SIMPLS-family estimator; an eligible bounded-block route is approximate rather than classical de Jong SIMPLS.

backend

Implementation backend: cpu, cuda, or metal. Multiple values are tuned in the inner loop. Metal requires float32 input and Apple Metal. R validates the request and constructs a reproducible grouped-fold plan. CPU and Metal use the compiled nested-CV coordinator for fold loops, fitting, prediction, and metric accumulation when selection uses accuracy, balanced accuracy or macro recall, Q2Y, or regression RMSD. Other selection metrics use R-level outer coordination around compiled single-CV and fitting kernels. CUDA uses the R coordinator around CUDA-native single-CV and outer-fit kernels; supported CUDA fits do not substitute a CPU estimator. Metal uses the same fixed operation split as pls(). When omitted, options(backend = ...) defines the session default. Every requested backend must be available; unavailable accelerator entries raise an error instead of using CPU.

n.cores

Number of CPU cores requested for supported BLAS/OpenMP host operations. An explicit value takes precedence over options(n.cores = ...). CUDA and Metal device parallelism is managed by their native runtimes; n.cores applies only to their host-side stages.

seed

Random seed used for outer/inner fold assignment and randomized SVD steps.

perm.test

Run a nested-CV permutation test. Independent rows are permuted individually. With repeated constrain values, complete blocks are exchanged only among groups with the same number of rows, preserving group sizes, within-group response structure, and class frequencies. Outer/inner folds and randomized-solver seeds are fixed across observed and null fits. The statistic is the median outer-CV value of the metric used for inner model selection.

times

Number of requested permutations. The corrected Monte Carlo p-value is (b + 1) / (B + 1), using the upper tail for metrics where larger is better and the lower tail for losses such as RMSD. B counts successful null fits only; failed fits are reported. A grouped test requires at least two constraint groups of equal size.

runn

Number of repeated runs.

kfold_inner

Inner-fold count, or "loocv" to leave out one constraint group at a time inside each outer training set.

kfold_outer

Outer-fold count, or "loocv" to leave out one constraint group at a time in the outer loop. In both loops, samples sharing the same constraint value are never split across training and test. A numeric fold count at least as large as the available number of groups has the same leave-one-group-out interpretation.

north

Number of orthogonal components removed by OPLS. The predictive count ncomp must fit within the predictor rank remaining after filtering. Direct PLS-LDA fits retain larger requested positions by repeating the last estimable result; other OPLS routes reject an unavailable predictive direction. Orthogonal components are not counted as predictive components.

kernel

Kernel type for kernel PLS: linear, rbf, or poly.

gamma

Kernel scale. Defaults internally to 1 / ncol(Xdata). For method = "kernelpls", multiple values are tuned in the inner loop.

degree

Polynomial kernel degree.

coef0

Polynomial kernel offset.

classifier

Classification decision rule. The default lda fits a regularized linear discriminant analysis classifier on the PLS latent scores. argmax selects the class with the largest PLS-DA response score.

bycol

For matrix-valued regression responses, calculate response-wise metrics in the returned metrics list. The default FALSE returns only aggregate metrics.

selection

Metric used by inner CV and by the permutation test. "auto" uses accuracy for classification and RMSD for regression. Classification also supports "balanced_accuracy", "AUROC", "lift_accuracy", "macro_precision", "macro_recall", "macro_f1", "kappa", "R2Y", and "Q2Y". Regression also supports "R2Y", "Q2Y", "RMSD", "MAE", "MAPE_percent", "RPD", "Pearson_r", and "Spearman_r". R2Y uses fitted responses for inner selection and is averaged across the selected outer training fits for the reported endpoint; as a training criterion, it can favor more complex models. Q2Y and the other predictive criteria use held-out predictions. Q2Y uses fold-training response means; the remaining metrics use the aggregate definitions in evaluate(). Binary AUROC selection pools the continuous inner-fold class scores before calculating the criterion; the second factor level is treated as positive. RMSD, MAE, and MAPE_percent are minimized; all other criteria are maximized. R2Y and Q2Y for classification operate on dummy-coded responses. Incompatible task/metric combinations raise an error before fitting. The former names "r2" and "q2" are rejected as ambiguous.

return_splits

Return a character matrix named split_index when TRUE. Rows use the original sample positions. Outer columns contain "training" or "test"; inner columns additionally use "outer_test" for samples excluded by the corresponding outer fold. Column names encode the run and outer/inner fold. The default FALSE avoids allocating this additional matrix.

...

Optional SVD tuning controls forwarded to the selected backend. Use the same compact names documented in fastsvd(), such as oversample and power. Vector values are tuned in the inner loop.

Details

For LDA classification, each inner and outer training fold uses only the classes represented in that fold and maps predictions back to the original factor levels. A class absent from a fold's training data cannot be predicted in that fold; its held-out observations remain in the reported metrics. With CUDA, nested outer/inner orchestration remains in R, while each supported single-CV and outer fit uses its native CUDA route. Unsupported accelerator requests fail explicitly and never fall back to CPU. In either task, a fold that estimates fewer latent components than requested uses its available component prefix, and higher requested prefixes repeat its last estimable prediction and metric. A zero-component regression fold predicts its training-response mean. A zero-component LDA fold uses empirical training-class priors and finite log-prior scores; a single-class training fold predicts its represented class. Inner tuning paths retain every requested component and record fold-level effective counts. Ties are resolved in favor of the earliest, and therefore lowest, requested component count.

Value

A list with the following elements. metrics$cross_validated contains one complete evaluate() result per repeated outer-CV run, and metrics$aggregate evaluates the final vote-aggregated or averaged prediction. For one run, its fold-aware Q2 is also attached to aggregate; after several runs, no unique fold-training reference exists for the averaged prediction, so aggregate Q2 is NA. metrics$definitions records the exact R2Y and Q2Y denominator conventions.

  • results: list with one element per repeated run. Each run stores Ypred/pred, the outer fold assignment, best_ncomp selected in each outer fold, fold-level best_parameters, compact inner-CV metric summaries in inner, selected outer-fold effective_ncomp, run-level metric_name and metric_value, and the default backend and method. Each inner summary contains its fold-by-prefix effective_ncomp matrix.

  • Ypred: final cross-validated predictions. For classification, repeated runs are combined by voting; for regression, numeric predictions are averaged across runs.

  • Q2Y: one outer cross-validated Q2 value per repeated run. Numeric responses use each outer fold's training-response mean. For factor responses this is the mean outer-fold dummy-response PLS-DA Q2 using outer-training class proportions and is not classification accuracy.

  • R2Y: one training-fit R2 value per repeated run, averaged across the selected outer-fold models.

  • RMSD: one held-out RMSD value per repeated run for numeric responses; NA for classification.

  • metric_name: metric used for each repeated run. It is held-out unless selection = "R2Y" explicitly requests the fitted-response criterion.

  • bcomp: most frequently selected component count across outer folds and repeated runs.

  • backend, method: default backend and PLS method supplied to the call. If vector-valued methods or backends are tuned, selected fold-level values are stored in results[[run]]$best_parameters.

  • selection_metric: criterion used by the inner CV loop.

  • degenerate_inner_folds, constant_classifier_fallback, minimum_positive_training_count, and minimum_negative_training_count: audit information for single-class inner-training partitions.

  • component_selection_informative and component_selection_informative_by_outer_fold: whether discrimination could inform component selection globally and in each outer fold.

  • outer_discrimination_estimable and non_estimable_outer_folds: flags for outer-training partitions containing both classes. A single-class outer-training partition returns its constant class rather than terminating the run.

  • acc_tot: classification-only text summary of correctly classified samples and percentage accuracy.

  • conf: classification-only confusion matrix printed as counts and column percentages.

  • vote_counts: classification-only vote-count matrix with one row per sample and one column per class.

  • accuracy: classification-only decoded-label accuracy, one value per repeated run.

  • balanced_accuracy: classification-only unweighted mean of class recalls, one value per repeated run.

  • medianR2Y, CI95R2Y, medianQ2Y, CI95Q2Y, medianRMSD, CI95RMSD: repeated-run summaries returned only when runn > 1.

  • permutation_metric, permutation_observed, and permutation_sampled: metric name, observed median, and permuted medians returned when perm.test = TRUE. Q2Ysampled is retained only when Q2 is the selected permutation metric.

  • p.value: permutation-test p-value returned when perm.test = TRUE.

  • permutation_unit, permutation_group_sizes_preserved, permutation_class_frequencies_preserved, permutation_folds, permutation_solver_seed, permutation_requested, permutation_completed, permutation_failed, and permutation_errors: the exchangeability contract and null-fit audit.

  • split_index: sample-by-split membership matrix for every outer and inner fold, returned only when return_splits = TRUE.

Examples

idx <- c(seq_len(10), 51:60, 101:110)
X <- as.matrix(iris[idx, seq_len(4)])
y <- factor(iris[idx, 5])
dcv <- pls.double.cv(X, y,
    ncomp = seq_len(2), runn = 1, kfold_inner = 2,
    kfold_outer = 2, method = "simpls", backend = "cpu", seed = 1
)
names(dcv)

fastPLS documentation built on Sept. 29, 2026, 1:06 a.m.