cf_glm_hv: Holdout validation for coarse-to-fine spatial generalized...

View source: R/cf_glm_hv.R

cf_glm_hvR Documentation

Holdout validation for coarse-to-fine spatial generalized linear mixed models (CF-GLMMs)

Description

Trains CF-GLMMs and selects the number of spatial scales through sequential holdout validation.

Usage

cf_glm_hv(
  y,
  x = NULL,
  coords,
  offset = NULL,
  train_rat = 0.75,
  id_train = NULL,
  alpha = 0.9,
  kernel = "exp",
  family = gaussian(),
  seed = 1234
)

Arguments

y

Vector of response variables (N x 1) including continuous, count, and binary responses, following an exponential family distribution.

x

Matrix of covariates (N x K).

coords

Matrix of 2-dimensional point coordinates (N x 2).

offset

Optional. Vector of offset variables (N x 1) included in the linear predictor, consistent with glm.

train_rat

Training sample ratio (default: 0.75). For small to moderate samples (N <= 30000), samples closest to the k-means centers are used for validation samples to stabilize training. For larger samples, training samples are drawn at random.

id_train

Optional. ID indicating training samples. If specified, the corresponding samples are used as training samples. Otherwise, training samples are chosen based on 'train_rat'.

alpha

Decay ratio of the kernel bandwidth in the coarse-to-fine training (default: 0.9). Values closer to one make the optimization more stringent but increase computation time.

kernel

Kernel type for modeling spatial dependence. '"exp"' for the exponential kernel (default) and '"gau"' for the Gaussian kernel.

family

Error distribution and link function specification, consistent with the 'family' argument of glm. Negative binomial responses: negbin() estimates the dispersion \theta (re-estimated on the training samples after each accepted scale); negbin(theta) or MASS::negative.binomial(theta) keeps it fixed. poisson(link = "identity") is supported with the mean floored at a small positive value.

seed

Random seed used for the training/validation split when 'id_train' is not supplied. Default is '1234'. Set to 'NULL' to allow a different split at each call (useful for assessing split sensitivity).

Value

A list with the following elements:

loss_hv

Final deviance loss for validation samples.

loss_hv_all

Deviance losses obtained at each learning step.

id_train

ID of training samples.

other

Other internally used output objects.

Author(s)

Daisuke Murakami

References

Murakami, D., Comber, A., Yoshida, T., Tsutsumida, N., Brunsdon, C., & Nakaya, T. (2025). Coarse-to-fine spatial GLMMs for scalable prediction and multiscale analysis. *ArXiv preprint*, 2605.01157. https://doi.org/10.48550/arXiv.2605.01157

See Also

cf_glm


spCF documentation built on Oct. 5, 2026, 5:07 p.m.