View source: R/item_restscore.R
| RMitemRestscore | R Documentation |
Computes observed and model-expected item-restscore correlations using
iarm::item_restscore(), and enriches the output with the absolute
difference between observed and expected values, item average locations, and
item locations relative to the sample mean person location.
RMitemRestscore(
data,
cutoff = NULL,
p_value = NULL,
correction = c("fwer", "fdr_bh", "fdr_by", "none"),
alpha = 0.05,
output = "kable",
sort,
p_adj = "BH"
)
data |
A data.frame or matrix of item responses. Items must be scored
starting at 0 (non-negative integers). Missing values ( |
cutoff |
Optional. Default
With a |
p_value |
Logical or
|
correction |
Character. Multiple-comparison correction applied across
items when |
alpha |
Numeric in (0, 1). Significance level for the |
output |
Character string controlling the return value. Either
|
sort |
Optional character string. When |
p_adj |
Character string specifying the p-value adjustment method
passed to |
Item-restscore correlations using Goodman-Kruskal's gamma (Kreiner, 2011) measure the association between a person's score on a single item and their total score on the remaining items (the "restscore"). Under a correctly fitting Rasch model, observed and model-expected correlations should agree closely.
Item parameters are estimated by conditional maximum likelihood via
psychotools::pcmodel() (a dichotomous item is a 2-category PCM); the
item-restscore statistic itself comes from iarm::item_restscore() and is
conditional on the total score, so it is invariant to the estimation engine.
Per-item average locations are the means of the CML thresholds, and the
person-location reference is the mean of the Warm WLE estimates.
Relative item location is defined as the item's average location minus the sample mean person location, providing a measure of item targeting.
The iarm package must be installed (it is in Suggests, not Imports).
The asymptotic p-value is miscalibrated. Without a cutoff, each
item is tested with (observed - expected) / SE against the standard
normal, where the SE is that of the observed gamma alone. The expected
gamma is estimated from the same data and correlates with the observed
one, so the SE is too large for the difference. The difference is also
biased upwards in small samples. Under a true Rasch model the test
therefore flags too many items as overfit in small samples and too few as
underfit at every sample size, and the adjustment chosen in p_adj does
not correct either. In simulation, 20 dichotomous items at n = 150 gave at
least one BH-flagged item in 13 percent of datasets (30 percent when 1.5
logits off target), almost all of them overfit, while 9 polytomous items
at n = 1000 gave about 1 percent. Pass the object from
RMitemRestscoreCutoff() as cutoff for p-values from a parametric
bootstrap null instead.
Bootstrap p-values. When p_value = TRUE, each item's observed
Difference is compared against its simulated null distribution (from
cutoff$results), studentised by the bootstrap mean and SD. Because the
bootstrap mean is subtracted, the small-sample upward bias of the
difference is removed, and because the bootstrap SD is used, the
correlation between observed and expected gamma is accounted for. The
marginal p-value is the two-sided Monte-Carlo p-value
(1 + #{|t*| >= |t|}) / (B + 1), and correction = "fwer" applies the
Westfall-Young studentised-max step-down across items (Ferreira, 2024).
The two directions do not get equal shares of alpha. Gamma is bounded at
1, so the null of Difference is left-skewed and a two-sided test on
|t| rejects more often in the long (underfit) tail. In simulation under
a true Rasch model the total rate was nominal, but the per-item marginal
rate was about 3 percent for underfit and 2 percent for overfit with 20
dichotomous items at n = 150 and 1.5 logits off target, narrowing toward
equal shares with polytomous items and larger samples. With
correction = "fwer", over four such conditions, at least one item was
flagged as underfit in 3.7 percent of datasets and as overfit in 1.6
percent, 5.0 percent in total. An equal-tailed test, with alpha/2 for each
direction, was evaluated and not adopted: it detected 2 to 3 percentage
points fewer underfitting items and no more overfitting ones.
Reproducibility. The bootstrap p-values depend on the simulated
null, so two analyses of the same data with different seeds can disagree
about items near the decision boundary. In simulation, in conditions
chosen to include such items, two seeds disagreed about the flag of at
least one item in about 10 percent of analyses at 400 iterations and about
6 percent at 1000 with correction = "fwer", and in 13 and 11 percent with
correction = "fdr_bh". Set seed in RMitemRestscoreCutoff() for a
reproducible analysis, and use 1000 or more iterations for a final one.
The direction in Flagged follows the studentised value, not the sign of
Difference. At small sample sizes the null mean of Difference is
slightly positive, so an item can be flagged "underfit" while its
Difference is still just above zero. This is rare.
If output = "kable": a knitr_kable object (plain text table via
format = "pipe") with columns for item name, observed and expected
restscore correlations, the signed difference (observed minus
expected), adjusted p-value, the Flagged misfit label, and item
location relative to the sample mean person location.
If output = "dataframe": a data.frame with columns Item, Observed,
Expected, Difference, p_adjusted, Flagged, and
Relative_location. Flagged is "overfit" (observed above expected,
adj. p < .05), "underfit" (below, adj. p < .05), or "" (not flagged).
With a cutoff, p_adjusted is replaced by Diff_low and Diff_high,
followed by p_restscore (marginal two-sided bootstrap p-value) and
padj_restscore (corrected p-value) when p_value resolves to TRUE.
Flagged then reflects padj_restscore < alpha, or the interval when
p_value = FALSE.
The Difference column is signed (observed minus expected):
positive values indicate that the item correlates more strongly with
the rest-score than the Rasch model predicts (over-discrimination /
overfit, often associated with local dependence), and negative
values indicate weaker-than-expected association (under-discrimination
/ underfit, often associated with multidimensionality or noise).
The marginal p-value controls the error rate of a single comparison: for
one item (or item pair) decided on in advance it is the relevant value. But
scanning all k comparisons and flagging whichever fall below alpha tests
k hypotheses at once, so the chance of at least one false flag inflates to
roughly 1 - (1 - \alpha)^k (e.g. about 34% for k = 8 at
alpha = 0.05) – even when every marginal p-value is correctly calibrated.
The corrected (adjusted) p-value controls this: correction = "fwer" bounds
the probability of any false flag (strict, lower power), while "fdr_bh" /
"fdr_by" bound the expected proportion of false flags among those raised
(a more lenient middle ground). Rule of thumb: use the marginal p-value for a
single pre-specified comparison, and a corrected p-value when screening the
whole table – the usual workflow.
Item infit and item-restscore both judge each item against expectations from a Rasch model fitted to all items, so misfitting items change what is expected of the others. A flag in one direction can therefore produce flags in the other. Overfitting items can push fitting items into the underfit range, and underfitting items, for example from a second dimension, can make fitting items look overfit. How large the effect is depends on how strong the misfit is and how many items share it. In simulation with seven polytomous items, two overfitting items led to about 6 percent of the fitting items being flagged as underfit, and one or two mildly multidimensional items to 1 to 2 percent being flagged as overfit. When items are flagged in both directions, consider removing the clearest misfit and testing again before interpreting the rest.
Kreiner, S. (2011). A Note on Item–Restscore Association in Rasch Models. Applied Psychological Measurement, 35(7), 557–561. \Sexpr[results=rd]{tools:::Rd_expr_doi("10.1177/0146621611410227")}
Ferreira, J. A. (2024). Methods of testing a 'small' or 'moderate' number of hypotheses simultaneously. Journal of Statistical Theory and Practice, 19(6). \Sexpr[results=rd]{tools:::Rd_expr_doi("10.1007/s42519-024-00412-4")}
RMitemRestscoreCutoff,
RMitemRestscorePlot
if (requireNamespace("iarm", quietly = TRUE)) {
# Simulate binary item response data (8 items, 200 persons)
set.seed(42)
sim_data <- as.data.frame(
matrix(sample(0:1, 200 * 8, replace = TRUE), nrow = 200, ncol = 8)
)
colnames(sim_data) <- paste0("Item", 1:8)
# Default kable output
RMitemRestscore(sim_data)
# Sorted by absolute difference
RMitemRestscore(sim_data, sort = "diff")
# Return as data.frame for further processing
df <- RMitemRestscore(sim_data, output = "dataframe")
# Bootstrap null distribution, flagging on Westfall-Young corrected
# p-values (use 1000 or more iterations in a final analysis)
if (requireNamespace("ggdist", quietly = TRUE)) {
cutoff_res <- RMitemRestscoreCutoff(sim_data, iterations = 100,
parallel = FALSE, seed = 42)
RMitemRestscore(sim_data, cutoff = cutoff_res)
}
}
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.