View source: R/est_LM_cat_multinom.R
| est_LM_cat_multinom | R Documentation |
Function that estimates a hidden Markov model to a univariate categorical longitudinal response.
The observed response is generated by an unobserved first-order Markov chain
with k latent states. Conditional on the latent state, responses are
independent over time. Covariates may affect both the initial latent-state
probabilities and the transition probabilities, and linear equality
constraints may be imposed separately on transition intercepts and covariate
effects.
est_LM_cat_multinom(resp_name, formula, k, data, reltol = 10^-10,
rand.start = FALSE,
nstart = 1, baseline = c("initial", "central"),
model_int = c("all", "const", "dist1", "dist2",
"dist3", "two", "symm", "rsymm", "diff"),
model_cov = c("all", "const", "dist1", "dist2", "dist3",
"two", "symm", "rsymm", "diff"),
formula_init = NULL,
par.init = list(la = NULL, psi = NULL, Pr = NULL),
output = FALSE, fort = TRUE)
resp_name |
Character string giving the name of the observed response variable in
|
formula |
A one-sided formula specifying the covariates in the transition model,
for example |
k |
Positive integer giving the number of latent states. A one-state model is
allowed; in that case only the response probabilities and fit criteria are
estimated, and the arguments |
data |
Data frame in long format. The first column is used as the subject identifier and the second as the measurement occasion, irrespective of their names. Subject identifiers must be consecutive integers, records must be ordered by subject and occasion, and occasions for each subject must be consecutive and start at 1. Individual trajectories may have different lengths. Rows with missing values in the response or model variables should be removed before fitting. |
reltol |
Positive numeric value giving the relative log-likelihood tolerance used
to stop the EM algorithm. Iteration stops when the relative increment in
log-likelihood is no greater than |
rand.start |
Logical value controlling initialization when |
nstart |
Positive integer giving the number of fits from different starting values. If greater than 1, the first fit uses the k-means initialization and the remaining fits use random initial allocations; the fit with the largest log-likelihood is returned. |
baseline |
Character string specifying the reference category in the multinomial
logit parameterization for the transition probabilities. With |
model_int |
Character string specifying the equality-constraint structure for the
transition-logit intercepts. Available choices are |
model_cov |
Character string specifying the equality-constraint structure for the
transition-logit regression coefficients. The available choices are the
same as for |
formula_init |
Optional one-sided formula specifying the covariates for the initial
latent-state probabilities. Only measurements at occasion 1 are used. If
|
par.init |
Named list of optional starting values. Component |
output |
Logical value. If |
fort |
Logical value selecting the numerical implementation. If |
Let U denote the latent state for subject i at occasion
t. The measurement model is
P(Y = y \mid U = u), represented by the returned
matrix Pr.
The arguments model_int and model_cov apply the same family of
constraint structures to different parameter blocks: model_int acts on
the transition intercepts, whereas model_cov acts componentwise on the
transition covariate coefficients. The two structures can be selected
independently. With baseline = "central", let a_{uv} denote a
transition parameter from origin state u to destination state v,
with v \ne u. The available structures are:
"all"No equality constraints: every transition has its own parameter.
"const"A common parameter is used for all off-diagonal transitions. For
model_cov, each covariate therefore has one common effect across all
transitions.
"dist1"Parameters are grouped by signed displacement v-u. Transitions with
the same direction and the same size of movement share a parameter.
"dist2"Parameters are grouped by absolute displacement |v-u|; movements of
the same size share a parameter in both directions.
"dist3"Parameters depend on |v-u|, with opposite signs for upward and
downward movements of the same size.
"two"Two parameters are used: one for transitions with v>u and one for
transitions with v<u, irrespective of movement size.
"symm"Opposite transitions share a parameter, so that
a_{uv}=a_{vu}.
"rsymm"Opposite transitions have parameters with reversed signs, so that
a_{uv}=-a_{vu}.
"diff"Transition parameters are differences between destination- and
origin-state parameters, a_{uv}=a_v-a_u, with the state-1 parameter
fixed at zero for identifiability.
The distance-based structures "dist1", "dist2",
"dist3", and "two" are intended for substantively ordered
latent states. In particular, model_cov = "all" is the unconstrained
transition-covariate model, model_cov = "dist1" pools effects by signed
displacement, and model_cov = "const" imposes one effect per covariate
across all transitions. The argument model_int operates analogously on
the transition intercepts.
When nstart > 1, the returned object is the solution with the largest
log-likelihood and component lkv records the log-likelihood obtained in
each fit.
An object of class "LM_cat_multinom" containing:
PrA k by c matrix of estimated conditional response
probabilities, with latent states in rows and observed categories in columns.
lkThe maximized observed-data log-likelihood.
npThe number of free model parameters.
aicThe Akaike information criterion.
bicThe Bayesian information criterion, using the number of subjects as the sample size.
callThe matched call used to fit the model.
For k > 1, the object also contains:
laEstimated regression coefficients for the initial latent-state probabilities.
psiEstimated reduced parameter vector for the constrained transition logits.
etaExpanded transition-logit parameter vector obtained by applying the design
matrix constraints to psi.
WArray of posterior latent-state probabilities produced by the final
E-step, with dimensions n by k by TT,
where n is the number of subjects.
CLMatrix of modal posterior latent-state classifications by subject and
occasion, with dimensions n by TT.
If output = TRUE and k > 1, it additionally contains:
PivMatrix of subject-specific fitted initial probabilities, with dimensions
n by k.
PIArray of subject- and occasion-specific fitted transition probabilities,
with dimensions k by k by n by TT, where
TT is the maximum number of occasions.
If nstart > 1, component lkv contains the log-likelihood from
each attempted start.
Francesco Bartolucci and Silvia Pandolfi, University of Perugia (Italy); Luca Brusa and Fulvia Pennoni, University of Milano-Bicocca (Italy).
Bartolucci, F., Pandolfi, S., and Pennoni, F. (2026). Parsimonious parametrizations of transition matrices of Markov chain and hidden Markov models. Annals of Operations Research, 363, 233–266.
design_matrices
## Not run:
# Load and prepare data
data("data_SRHS_long")
srhs <- data_SRHS_long[,c(2,1,3:7)]
# 1 = poor, ..., 5 = excellent
srhs$srhs1 <- 6L - as.integer(srhs$srhs)
srhs$gender <- factor(srhs$gender)
srhs$race <- factor(srhs$race)
srhs$education <- factor(srhs$education)
srhs$age_c <- srhs$age - mean(srhs$age)
srhs$id <- as.integer(factor(srhs$id))
srhs <- srhs[order(srhs$id, srhs$t), ]
# Reduce the sample size for fast example
set.seed(123)
all_ids <- unique(srhs$id)
n_example <- min(1000L, length(all_ids))
example_ids <- sample(all_ids, size = n_example, replace = FALSE)
srhs_example <- srhs[srhs$id %in% example_ids,,drop = FALSE]
# Consecutive identifiers required by the estimation function
srhs_example$id <- match(srhs_example$id, example_ids)
srhs_example <- srhs_example[order(srhs_example$id, srhs_example$t), ]
# Estimate a HM-ALL-COV: benchmark without constraints on covariate effects
hm_all <- est_LM_cat_multinom("srhs1", ~ gender + race + education + age_c,
k = 3, data = srhs_example, baseline = "central",
model_cov = "all", model_int = "all",
output = TRUE, fort = TRUE)
print(hm_all)
summary(hm_all)
# Model with a common covariate effect across transitions
hm_const <- est_LM_cat_multinom("srhs1", ~ gender + race + education + age_c,
k = 3, data = srhs_example, baseline = "central",
model_cov = "const", model_int = "all",
output = TRUE, fort = TRUE)
print(hm_const)
summary(hm_const)
## End(Not run)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.