knitr::opts_chunk$set( collapse = TRUE, comment = "#>", message = FALSE, warning = FALSE, fig.width = 7, fig.height = 5.5 ) library(idiographic) data(srl) vars <- c("efficacy", "value", "planning", "monitoring", "effort")
The autoregressive models estimated in this package — ordinary and graphical VAR, multilevel VAR, unified SEM, GIMME, and their rolling variants — share a common set of assumptions about the input series. Each presumes weak stationarity, meaning that the mean, variance, and autocovariance of every series are constant across the observation window; linear, first-order (lag-one) dynamics; measurement occasions that are correctly ordered and approximately equally spaced; and non-degenerate variance in each variable [@bringmann2013; @epskamp2018mlvar]. Intensive longitudinal data seldom satisfy these conditions automatically. Series drift or trend across a protocol, their variance differs between its early and late phases, occasions are missed or unevenly attended, and an item is occasionally answered identically throughout. None of these violations is detected by the estimators themselves: a trending series, for instance, is fitted without error and its trend is absorbed into an inflated autoregressive coefficient, so the resulting network misrepresents the dynamics rather than failing outright.
The purpose of preprocess() is to make these assumptions inspectable, and
where necessary correctable, before a model is estimated and interpreted. It
constructs the same lag-one design that the estimators use — applying the same
variable selection, ordering, scaling, and within-person centring — and
evaluates each subject-series against the stationarity assumptions, returning a
set of diagnostic flags rather than a fitted network. It is not a required step:
the estimators accept the raw data and standardize and centre it internally. Its
role is diagnostic, and, where an assumption is violated, corrective, through
the detrend argument described below.
The estimators expect long format: one row per person-occasion, an id column, an
ordering column, and numeric time-varying indicators. The bundled srl data
hold self-regulated-learning indicators for 36 students measured over 156
occasions each; name is the student and day orders the occasions. This
vignette works with five indicators: efficacy, value, planning,
monitoring, and effort.
A temporal edge is only defined once the series is ordered: from -> to means
from at occasion $t-1$ predicts to at occasion $t$. Losing or reordering
lagged pairs changes that estimand, which is why the design is worth inspecting
before fitting.
A call to preprocess() takes the variables and the identifier column. Printing
the returned object reports the retained design and, where an assumption is at
risk, an explicit recommendation for addressing it.
pp <- preprocess(srl, vars = vars, id = "name") pp
The five indicators over 36 students give 5616 ordered rows. Of these, 5548 survive as complete current/lagged pairs: each student loses the first occasion to the initial lag, and a few students lose a little more to missing values. No series is constant, none shows near-unit-root persistence, and the unit-root screen is clear — but ten subject-series trip the linear-trend flag, and the message names the fix.
The summary() method aggregates the per-subject diagnostics to one row per
variable, reporting its mean dispersion and the number of subject-series that
trip each flag.
summary(pp)
The trend flags are concentrated in planning (5 of 36 students) and value
(2), with one each elsewhere. Across the flagged students planning drifts by
roughly 0.003 standardized units per occasion — small per step, but enough over
156 occasions to bias its lag-one coefficients. The n_high_ar and
n_unit_root columns are zero throughout: the non-stationarity here is a gentle
deterministic trend, not a random walk.
The counts table gives the lag-pair accounting per student, which is the
denominator behind every temporal edge.
head(pp$counts)
Each 156-occasion series offers 155 possible lag pairs; students with complete data retain all 155, while a student with a few missing occasions (here, Bob) retains fewer.
Each row of pp$diagnostics is one subject-series, and each flag marks a
different way stationarity can fail:
flag_trend — a significant linear slope over time; the mean is not
constant.flag_high_ar — a lag-one autocorrelation at or above 0.95; near-unit-root
persistence.flag_unit_root — an ADF-style screen suggests the series may not revert
to a stable mean.flag_mean_shift / flag_sd_shift — the first and second halves of the
series differ in mean or in spread; drift or changing volatility.flag_zero_variance — a constant series, which no network can use.flag_stationarity_risk is TRUE when any of these fire, and is the column the
summary() roll-up counts as n_flagged.
detrendpreprocess() standardizes and centres, but it does not remove trends on its
own — detrending changes what a temporal edge means, so it is a modelling
decision the analyst should make deliberately. The detrend
argument makes that decision explicit:
"none" (default) — diagnose only."auto" — transform only the flagged series: difference a stochastic trend
(unit root / near-unit-root), linearly detrend a deterministic trend, and
leave every stationary series untouched."linear" / "difference" — force one transform on every series."auto" is the option the diagnostic message recommended. It runs over all
subject-series and cleans only those that need it — no subsetting, one call.
clean <- preprocess(srl, vars = vars, id = "name", detrend = "auto") summary(clean)
The ten trend-flagged series are linearly detrended and the other 170 are left
as they were, so every n_trend count drops to zero. The retained-pair count is
unchanged because no series here needed differencing (which would cost the first
occasion of each block). One planning flag survives: it is a variance shift,
not a trend, and "auto" deliberately does not touch heteroscedasticity — that
is a substantive issue to weigh, not something to difference away.
When "auto"'s per-series choices are not what you want, detrend also takes a
named vector to set the method one variable at a time; unlisted variables
are left untouched. And checks restricts the screening to the risks you care
about — for example, flag only trends and unit roots and ignore variance drift.
clean2 <- preprocess(srl, vars = vars, id = "name", detrend = c(planning = "difference", value = "linear"), checks = c("trend", "unit_root")) summary(clean2)
The transformed lag-pair design is available in clean$pairs and its current
and lagged matrices in matrices(clean). The model-fitting functions take long
data rather than a preprocess_result, so an applied analysis must make the
same documented transformation to its long input before fitting; the audit does
not silently alter srl or any later model call.
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.