| RMreliability | R Documentation |
Computes three reliability indices for a Rasch / partial credit model:
the Person Separation Index (PSI) – WLE-based separation reliability – the
marginal reliability (native, CML test information integrated over the
estimated normal latent density), and Relative Measurement Uncertainty (RMU)
via RMUreliability() applied to plausible values from mirt::fscores().
RMreliability(
data,
conf_int = 0.95,
draws = 1000,
rmu_iter = 50,
estim = "WLE",
boot = FALSE,
boot_iter = 200,
parallel = TRUE,
n_cores = NULL,
seed = NULL,
verbose = FALSE,
theta_range = c(-10, 10),
output = "kable"
)
data |
A data.frame or matrix of item responses. Items must be scored starting at 0 (non-negative integers). |
conf_int |
Numeric in (0, 1). HDCI width for both bootstrap CIs and
RMU. Default |
draws |
Integer. Number of plausible-value draws drawn from the mirt
model for the RMU calculation. Default |
rmu_iter |
Integer. Number of times |
estim |
Character. Theta estimator used by |
boot |
Logical. If |
boot_iter |
Integer. Number of bootstrap iterations when
|
parallel |
Logical. Use parallel processing via |
n_cores |
Integer or |
seed |
Integer or |
verbose |
Logical. Print progress messages and a progress bar for
the bootstrap. Default |
theta_range |
Numeric length-2 vector. Theta limits passed to
|
output |
Character. |
Confidence intervals for PSI and Marginal reliability are obtained
by non-parametric bootstrap (resampling respondents; all three indices are
recomputed natively per resample, no model is refitted by mirt). The RMU
interval is the HDCI of correlations across plausible-value draws, averaged
over rmu_iter random splits of the draws.
Marginal reliability is the latent-density-weighted mean of the conditional
reliability curve, \int \sigma^2/(\sigma^2 + 1/I(\theta))\,
g(\theta)\,d\theta, where the test information I(\theta) is summed from
the CML item parameters and g is the estimated normal latent density
N(0, \sigma^2) (\sigma from marginal ML). Integrating over the
estimated latent variance, rather than the N(0,1) assumed by
mirt::marginal_rxx(), keeps it correct on the Rasch logit scale, where
\sigma is typically well above 1 and the N(0,1) assumption
underestimates reliability.
It is the model-based complement to the sample-based PSI, and the two are
now the same coefficient by two routes: PSI divides by the observed spread
of the WLE estimates, marginal reliability by the fitted latent density. A
large gap between them therefore does flag an off-target or non-normal
sample. Through version 1.2.0 this row used Green's subtractive
1 - \overline{1/I(\theta)}/\sigma^2 instead, under which much of the
PSI-to-marginal gap was an artefact of the differing formulas rather than a
property of the sample. See dev/TODO-reliability-form.md.
PSI is the WLE-based separation reliability,
1 - \overline{SEM^2} / \mathrm{Var}(\hat\theta), computed from CML item
thresholds (psychotools) and Warm's WLE person locations / analytic SEMs.
Respondents with extreme (min/max) raw scores are excluded – their boundary
estimates would inflate the person variance and overstate reliability.
(Earlier versions used eRm::SepRel() with MLE; the values can differ, most
noticeably for scales with many extreme scorers, e.g. dichotomous items.)
RMU is from Bignardi, Kievit, & Bürkner (2025), modified here to use mirt plausible values rather than fully Bayesian posterior draws (see Mislevy, 1991, for the plausible-values framework).
Marginal reliability here is the subtractive Green/Lord coefficient. Milanzi
et al. (2015) show that this form can fall below zero when the average error
variance exceeds the trait variance, which happens with few items or a
narrow sample, and it is floored at 0 above. RMreliabilityCurve() reports
the same quantity in the bounded ratio form alongside this one, so the two
can be compared directly. See dev/TODO-reliability-form.md for the open
question of which form this row should use.
Bootstrap iterations that fail to converge are silently dropped.
If output = "kable": a knitr_kable object with one row per metric.
If output = "dataframe": a data.frame with columns metric,
estimate, lower, upper, notes.
Bignardi, G., Kievit, R., & Bürkner, P. C. (2025). A general method for estimating reliability using Bayesian Measurement Uncertainty. PsyArXiv. \Sexpr[results=rd]{tools:::Rd_expr_doi("10.31234/osf.io/h54k8_v1")}
Green, B. F., Bock, R. D., Humphreys, L. G., Linn, R. L., & Reckase, M. D. (1984). Technical Guidelines for Assessing Computerized Adaptive Tests. Journal of Educational Measurement, 21(4), 347–360. \Sexpr[results=rd]{tools:::Rd_expr_doi("10.1111/j.1745-3984.1984.tb01039.x")}
Mislevy, R. J. (1991). Randomization-Based Inference about Latent Variables from Complex Samples. Psychometrika, 56(2), 177-196. \Sexpr[results=rd]{tools:::Rd_expr_doi("10.1007/BF02294457")}
Adams, R. J. (2005). Reliability as a measurement design effect. Studies in Educational Evaluation, 31(2), 162-172. \Sexpr[results=rd]{tools:::Rd_expr_doi("10.1016/j.stueduc.2005.05.008")}
Milanzi, E., Molenberghs, G., Alonso, A., Verbeke, G., & De Boeck, P. (2015). Reliability measures in item response theory: Manifest versus latent correlation functions. British Journal of Mathematical and Statistical Psychology, 68(1), 43-64. \Sexpr[results=rd]{tools:::Rd_expr_doi("10.1111/bmsp.12033")}
RMUreliability(), RMreliabilityCurve()
if (requireNamespace("ggdist", quietly = TRUE) &&
requireNamespace("eRm", quietly = TRUE)) {
set.seed(1)
RMreliability(eRm::raschdat1[, 1:20], draws = 1000)
# Bootstrap CI for PSI and Marginal
# (use more bootstrap iterations, e.g. 200+, in real analyses)
RMreliability(eRm::raschdat1[, 1:20], draws = 1000,
boot = TRUE, boot_iter = 25, parallel = FALSE, seed = 42)
}
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.