cf_dglm_hv: Holdout validation for coarse-to-fine dynamic (space-time)...

View source: R/cf_dglm_hv.R

cf_dglm_hvR Documentation

Holdout validation for coarse-to-fine dynamic (space-time) spatial GLMMs

Description

Trains a coarse-to-fine dynamic spatial GLMM (CF-DGLMM) and selects the spatial scales of a separable space-time cascade through progressive holdout validation. The companion cf_dglm refits the selected structure on the full sample and predicts. The model decomposes the link-scale linear predictor as g(\mu_{i,t}) = x_{i,t}'\beta + \sum_k f_k(s_i,t) + offset, where each scale-k field f_k is a per-knot AR(1) Kalman smoother in time combined with kernel kriging in space.

Usage

cf_dglm_hv(
  y,
  x = NULL,
  coords,
  time,
  offset = NULL,
  train_rat = 0.75,
  id_train = NULL,
  alpha = 0.9,
  kernel = "exp",
  family = gaussian(),
  rho = NULL,
  Q = NULL,
  tvc = NULL,
  q_tvc = NULL,
  seed = 1234
)

Arguments

y

Vector of response variables (N x 1) including continuous, count, and binary responses, following an exponential family distribution.

x

Matrix of covariates (N x K).

coords

Matrix of 2-dimensional point coordinates (N x 2). Rows sharing the same coordinates are treated as repeated observations of one location across time. The space-time panel may be unbalanced: the set of observed locations is allowed to differ from one time point to another (knots are placed on the union of locations and the per-knot AR(1) smoother bridges time points at which a knot has no nearby observation).

time

Vector of time indices (N x 1) identifying the time point of each observation. Any sortable type (integer, numeric, Date) is accepted.

offset

Optional. Vector of offset variable (N x 1) to be included in the linear predictor, consistent with glm.

train_rat

Training sample ratio (default: 0.75). Holdout is performed at the location level: a subset of locations (and all their time points) is held out for validation.

id_train

Optional. If specified, the corresponding samples are used as training samples; otherwise locations are chosen based on train_rat.

alpha

Decay ratio of the kernel bandwidth in the coarse-to-fine training (default: 0.9).

kernel

Kernel type for spatial dependence. "exp" for the exponential kernel (default) and "gau" for the Gaussian kernel.

family

Error distribution and link function, consistent with the family argument of glm. Functionality has been confirmed for gaussian(), poisson(), and binomial(). Negative binomial responses: negbin() estimates the dispersion \theta (re-estimated on the training samples after each accepted scale); negbin(theta) or MASS::negative.binomial(theta) keeps it fixed. poisson(link = "identity") is supported with the mean floored at a small positive value.

rho, Q

Optional AR(1) temporal parameters (autocorrelation and innovation variance). When NULL (default) a single global (rho, Q) is estimated by maximum marginal likelihood.

tvc

Optional. Covariates whose regression coefficients are allowed to vary over time, given as covariate names or as integer column indices into x. The remaining coefficients are constant. The intercept is always kept constant (a time-varying intercept is confounded with the temporal mean of the spatial field). NULL (default) keeps all coefficients constant.

q_tvc

Optional. Innovation (drift) variance of the random walk followed by the time-varying coefficients. When NULL (default) it is estimated from the data.

seed

Random seed for the training/validation split and knot placement (default 1234). Set to NULL for a random split.

Value

A list of class "cf_dglm_hv" with the following elements:

loss_hv

Holdout deviance of the selected model, evaluated at the validation locations. Fits of the same data share the same split, so this value compares models directly, whatever number of scales each selected.

loss_hv_all

The validation loss after every learning step.

e_summary

Out-of-sample accuracy at the validation locations of the model trained on the training locations only: deviance-based pseudo R-squared (validation_Pseudo-R2, ordinary R-squared in the Gaussian case), validation_RMSE and validation_MAE. Unlike the e_summary of cf_dglm, which scores the full-sample refit at those same points, this one never saw them.

val_pred

The validation predictions behind e_summary: one row per held-out observation with its location index (loc), time point (time), observed response (y) and predicted mean (pred) on the response scale.

id_train

Row indices of the training observations.

other

Internal objects reused by cf_dglm.

call

The matched call.

Author(s)

Daisuke Murakami

References

Murakami, D. (2026). Fast covariance-free spatiotemporal modeling via coarse-to-fine learning. *ArXiv preprint*, 2608.03449.

See Also

cf_dglm, cf_glm_hv


spCF documentation built on Oct. 5, 2026, 5:07 p.m.