lavH1: The unrestricted (H1) model

View source: R/lav_lavh1.R

lavH1R Documentation

The unrestricted (H1) model

Description

Compute only the summary statistics of the unrestricted (H1) model: the saturated mean vector and variance-covariance matrix (and thresholds, if some variables are ordered) that normally end up in the @h1 slot of a fitted lavaan object. Unlike sem or cfa, no model is needed. If a model is provided anyway, it is only used to determine the relevant variables, their role (exogenous, ordered, per level, ...) and the data-related options, so that the result is exactly the h1 information a fit of that model would use.

If the data are incomplete and missing = "ml" (fiml), the summary statistics are the EM estimates of the saturated model; this is a convenient way to obtain fiml estimates of standard descriptive statistics. For multilevel data (using the cluster argument), the within- and between-level summary statistics (and the intraclass correlations) are computed.

Usage

lavH1(model = NULL, data = NULL, ordered = NULL, sampling_weights = NULL,
      group = NULL, cluster = NULL, missing = "default",
      ov_names_x = NULL, ov_names_l = list(),
      estimator = "default", se = "default", test = "none", ...,
      output = "list", wls_v = NULL, gamma = NULL, add_extra = TRUE,
      add_labels = TRUE, add_class = TRUE, drop_list_single_group = TRUE)

Arguments

model

Optional. Either a description of the user-specified model (as in sem), or a fitted lavaan object. If a model is provided, it is only used to determine the variables, their role, and the data-related options. For convenience, the first argument may also be the data itself (a data.frame), in which case the data argument should stay empty.

data

A data.frame containing the observed variables. If no model is provided, all variables in the data frame are used (except the variables given by the group, cluster and sampling_weights arguments).

ordered

Character vector declaring additional variables as ordered (ordinal). Variables that are declared as ordered factors in the data frame are always treated as ordinal. If TRUE, all variables are treated as ordinal.

sampling_weights

A variable name in the data frame containing sampling weight information.

group

A variable name in the data frame defining the groups in a multiple group analysis.

cluster

A variable name in the data frame defining the clusters in a two-level dataset.

missing

How to handle missing data. If "ml" (alias: "fiml"), an EM algorithm is used to compute the saturated (h1) summary statistics using all available information. If "two.stage" or "robust.two.stage" (continuous data only), the EM saturated summary statistics are returned together with the stage-1 nacov (the asymptotic covariance matrix of those statistics) when output = "lavMoments", so that a follow-up analysis based on the summary statistics only reproduces the same (two-stage) estimates, standard errors and test statistics as when the raw data were provided. See lavOptions for further options.

ov_names_x

Character vector. Variables that need to be treated as exogenous. Only used if no model is provided.

ov_names_l

List. Only used for two-level data when no model is provided: the variable names per level. By default, all variables are used at both levels.

estimator

The estimator (see lavOptions); the default corresponds to maximum likelihood ("ML") for continuous data, and to "WLSMV" when some variables are ordered. The estimator determines, among other things, how the summary statistics are computed (for example, pairwise maximum likelihood for the polychoric correlations when some variables are ordered), and whether the h1 loglikelihood is available.

se

Only used if output = "lavaan": the standard errors for the parameters (= the summary statistics) of the unrestricted model. By default, standard errors are computed (matching the estimator), so that summary() shows the standard errors, Wald (z) statistics and p-values of the summary statistics; use se = "none" to skip them. See lavOptions for possible options.

test

Only used if output = "lavaan". See lavOptions for possible options.

...

Optional parameters that are passed to the lavaan function (for example fixed_x, conditional_x, h1_missing_method, ...).

output

Character. If "list" (the default), a list with the h1 summary statistics per group (and per level in the two-level setting): the (residual) variance-covariance matrix ($cov), the mean vector ($mean, if a meanstructure was used), the thresholds ($th, if some variables are ordered), and the regression slopes of the exogenous covariates ($slopes, if conditional_x = TRUE). If "lavaan", the unrestricted model itself is returned as a fitted lavaan object. If "lavMoments" (alias "moments"), an object of class lavMoments: a list of summary statistics whose components are named after the corresponding arguments of lavaan: $sample_cov, $sample_nobs, $sample_mean (if a meanstructure was used), $sample_th (if some variables are ordered; including the th.idx attribute), $wls_v and $nacov (if requested, see below), and $lav_options (the options a refit from these summary statistics must use, such as conditional_x). For conditional_x = TRUE, the residual slopes and the exogenous-covariate moments travel as attributes of $sample_cov (following the sample_cov argument convention of lavaan). The object is structured so that it can be passed directly as the data= argument of lavaan/cfa/sem to (re)fit a model from summary statistics only. For categorical data, the wls_v (weight) and nacov (gamma) matrices are included by default, so that the refit reproduces the estimates, standard errors and test statistics exactly. (When reading a lavMoments object, the fitting functions also accept the dotted component names sample.cov, sample.nobs, WLS.V, NACOV and lavOptions, as produced by older code.)

wls_v

Logical. If TRUE, the wls_v weight matrix (the h1 expected information; full or diagonal, depending on the estimator) is also computed and returned. If NULL (the default), it is set to TRUE for categorical data when output = "lavMoments", and to FALSE otherwise.

gamma

Logical. If TRUE, the gamma (nacov) matrix (the asymptotic variance matrix of the summary statistics) is also computed and returned. If NULL (the default), it is set to TRUE for categorical data when output = "lavMoments", and to FALSE otherwise.

add_extra

Logical. If TRUE, the list output also contains the h1 loglikelihood ($logl, if available, and $logl_group for multiple groups), the number of observations ($nobs), the coverage proportions ($coverage, only if the missing data were handled), and the intraclass correlations ($icc, only in the two-level setting). Ignored when output = "lavMoments".

add_labels

Logical. If TRUE, variable names are attached to the returned statistics.

add_class

Logical. If TRUE, lavaan print classes are attached to the returned statistics.

drop_list_single_group

Logical. If TRUE, the per-group list structure is dropped when there is only a single group.

Value

If output = "list" (the default), a list with the h1 summary statistics (see the output argument above). If output = "lavaan", a fitted lavaan object for the unrestricted model; its summary() shows the standard errors, Wald (z) statistics and p-values of the summary statistics (for example, to flag significant correlations when correlation = TRUE). The object may also serve as an ingredient for two-stage estimation procedures. If output = "lavMoments", an object of class lavMoments that can be passed directly as the data= argument of lavaan/cfa/sem.

Reproducing a fit from an output = "lavMoments" object

A lavMoments object reproduces the original raw-data fit exactly for the estimators whose estimates, standard errors and test statistics are functions of the summary statistics only. This includes maximum likelihood ("ML") and generalized least squares ("GLS"); the robust variants ("MLM", "MLMV", "MLMVS", ... and the Satorra-Bentler / scaled-shifted / mean-variance-adjusted tests), provided the gamma (nacov) matrix is requested (gamma = TRUE); the diagonally weighted least squares estimators for ordered data ("WLSMV", "WLSM", "DWLS", ..., which include wls_v and nacov by default); and two-stage maximum likelihood for incomplete continuous data (missing = "two.stage" or "robust.two.stage").

For the estimators whose weight matrix is estimator-specific (for example the identity weight of the "ULS" family, or the ADF weight of continuous "WLS"), pass the SAME estimator to lavH1() as will be used in the refit, so that the matching wls_v is returned.

A raw-data fit cannot be reproduced from summary statistics when the standard errors or test statistic depend on the individual cases: the Huber-White ('robust') sandwich ("MLR", se = "robust.huber.white"), the first-order/outer-product information ("MLF"), cluster-robust (cluster =) standard errors, the bootstrap (se = "bootstrap"), pairwise maximum likelihood ("PML"), and direct full-information maximum likelihood for incomplete data (missing = "ml"; use missing = "two.stage" instead).

See Also

lavCor for (polychoric, polyserial and Pearson) correlations (which uses the same machinery), and lavInspect with what = "h1" to extract the h1 summary statistics from a fitted lavaan object.

Examples

# saturated summary statistics of the Holzinger & Swineford data
lavH1(HolzingerSwineford1939[, paste0("x", 1:9)])

# fiml (EM) estimates of the descriptive statistics with missing data
HS.missing <- HolzingerSwineford1939
HS.missing$x5[1:10] <- NA
lavH1(HS.missing[, paste0("x", 1:9)], missing = "ml")

# a model is accepted, but only determines the variables and options
HS.model <- ' visual  =~ x1 + x2 + x3
              textual =~ x4 + x5 + x6
              speed   =~ x7 + x8 + x9 '
h1 <- lavH1(HS.model, data = HolzingerSwineford1939, group = "school")

# the unrestricted model as a fitted lavaan object; summary() shows the
# standard errors, z-statistics and p-values of the summary statistics
fit.h1 <- lavH1(HolzingerSwineford1939[, paste0("x", 1:9)],
                output = "lavaan")
summary(fit.h1)

# which correlations are significant?
fit.cor <- lavH1(HolzingerSwineford1939[, paste0("x", 1:9)],
                 output = "lavaan", correlation = TRUE)
summary(fit.cor)

# categorical data: extract all summary statistics (including wls_v and
# nacov), then (re)fit a model from the summary statistics only
Data <- HolzingerSwineford1939[, paste0("x", 1:9)]
Data[] <- lapply(Data, cut, breaks = 3, labels = FALSE)
Data[] <- lapply(Data, ordered)
model <- ' visual  =~ x1 + x2 + x3
           textual =~ x4 + x5 + x6
           speed   =~ x7 + x8 + x9 '
moments <- lavH1(model, data = Data, output = "lavMoments")
moments
fit <- cfa(model, data = moments)   # identical to cfa(model, data = Data)

# two-stage ML for incomplete (continuous) data: the EM saturated moments
# plus the stage-1 nacov allow a moments-only refit to reproduce the exact
# two-stage estimates, standard errors and (scaled) test statistics
HS.incomplete <- HolzingerSwineford1939[, paste0("x", 1:9)]
HS.incomplete$x1[1:20] <- HS.incomplete$x5[15:40] <- NA
model2 <- ' visual =~ x1 + x2 + x3
            textual =~ x4 + x5 + x6 '
moments2 <- lavH1(model2, data = HS.incomplete, missing = "two.stage",
                  output = "lavMoments")
fit2 <- cfa(model2, data = moments2)
# same as: cfa(model2, data = HS.incomplete, missing = "two.stage")

lavaan documentation built on Oct. 8, 2026, 5:06 p.m.