hier_boot_jsd_model: Hierarchical bootstrap for JSD-based models

View source: R/hier_boot_jsd_model.R

hier_boot_jsd_modelR Documentation

Hierarchical bootstrap for JSD-based models

Description

Performs a hierarchical bootstrap: resample groups with replacement, resample tokens within each sampled group, compute JSD per group, fit a model to the bootstrap JSD values, and repeat.

Usage

hier_boot_jsd_model(
  data,
  group_col,
  category_col,
  features,
  formula,
  fit_fun = NULL,
  n_outer = 200,
  min_tokens = 20,
  eps = 1e-06,
  progress = TRUE,
  ...
)

Arguments

data

Data frame with at least: group_col, category_col, features, and any predictors used in the model.

group_col

String: grouping variable (e.g., "speaker").

category_col

String: category variable with 2 levels (e.g., "vowel").

features

Character vector of acoustic feature columns.

formula

Model formula to pass to fit_fun (e.g., jsd_beta ~ s(age) + s(region, bs = "re")).

fit_fun

A function that takes ⁠(formula, data, ...)⁠ and returns a fitted model. Defaults to mgcv::gam if available, otherwise stats::lm.

n_outer

Number of hierarchical bootstrap replicates.

min_tokens

Minimum within-group tokens required.

eps

Small epsilon for bounding JSD in (0, 1) if using Beta family.

progress

Logical; if TRUE, prints progress every 10 replicates.

...

Additional arguments passed to fit_fun.

Details

This lets you propagate measurement uncertainty in JSD into model parameters (e.g., GAM/LMM coefficients).

Value

A tibble with columns:

  • boot_id - bootstrap replicate index

  • term - model term

  • estimate - estimate for that term in that replicate

Examples

set.seed(2026)
speakers <- paste0("s", 1:4)
dat <- data.frame(
  speaker = rep(speakers, each = 60),
  age = rep(c(22, 35, 48, 61), each = 60),
  vowel = rep(rep(c("ih", "eh"), each = 30), 4)
)
dat$f1 <- rnorm(
  nrow(dat),
  mean = ifelse(dat$vowel == "ih", 500, 560) + dat$age * 0.3,
  sd = 55
)

hier_boot_jsd_model(
  data = dat,
  group_col = "speaker",
  category_col = "vowel",
  features = "f1",
  formula = jsd_beta ~ age,
  fit_fun = stats::lm,
  n_outer = 3,
  min_tokens = 20,
  progress = FALSE
)

phontrast documentation built on Oct. 7, 2026, 5:06 p.m.