| ctt | R Documentation |
Computes classical test theory (CTT) statistics from a scored item
response data set (dichotomous or polytomous). For each item, computes
its difficulty, an item-total correlation reflecting how well the item
discriminates between high- and low-scoring examinees, and Cronbach's
alpha recomputed with that item removed, with optional flagging of items
whose difficulty or discrimination falls outside conventional quality
thresholds. At the test level, computes Cronbach's alpha as a reliability
estimate (in both raw and standardized forms), the associated standard
error of measurement (SEM), and the average item difficulty and
discrimination across the test. The distribution of total (raw) scores
across examinees is also tabulated. The result is returned as an object of
class "ctt", with print.ctt() and summary.ctt() methods that display
a condensed and a full report, respectively, analogous to how
est_irt() pairs with print.est_irt()/summary.est_irt().
ctt(
data,
item.id = NULL,
cats = NULL,
correct = FALSE,
missing = NA,
flag = TRUE,
crit.p = c(0.1, 0.95),
crit.dis = 0.2
)
data |
A data frame or matrix of already-scored item responses, with
examinees in rows and items in columns. Item scores must range from 0 to
|
item.id |
A character vector of item identifiers, in the same order
as the columns of |
cats |
A numeric vector giving the number of score categories for
each item (e.g., 2 for a dichotomous item), following the |
correct |
Logical. Both the raw (uncorrected) item-total correlation -
where an item is correlated with the total score that includes its own
contribution - and the corrected item-total correlation - excluding its
own contribution - are always computed and reported as separate columns.
This argument only controls which of the two is used for flagging: if
|
missing |
A value indicating missing responses in |
flag |
Logical. If |
crit.p |
A numeric vector of length two giving the lower and upper
difficulty bounds used for flagging: difficulty below the first value is
flagged as too difficult, and difficulty above the second value is
flagged as too easy. Default is |
crit.dis |
A single numeric value giving the minimum acceptable
discrimination (item-total correlation); items strictly below this value
are flagged as poorly discriminating. Default is |
Difficulty for item j is the mean observed item score divided by the item's maximum possible score,
p_j = \frac{\bar{X}_j}{m_j},
where m_j is cats[j] - 1. For a dichotomous item (cats[j] = 2)
this reduces to the familiar proportion-correct difficulty index; for a
polytomous item it expresses the average score as a proportion of the
maximum attainable score, keeping difficulty on a common 0-1 scale
regardless of the number of score categories.
Discrimination for item j is the Pearson correlation between the item
score and the total score, always computed and reported in two forms: the
raw (uncorrected) item-total correlation, which correlates the item with
the total score that includes the item's own contribution, and the
corrected (item-excluded) item-total correlation, which excludes it. For a
dichotomous item, the raw item-total correlation is mathematically
equivalent to the point-biserial correlation. correct selects which of
the two feeds the discrimination flagging criterion (crit.dis); both are
always reported as separate columns regardless of correct (this mirrors
how, e.g., psych::alpha() reports the raw item-total correlation as
raw.r and the corrected version as r.drop).
Alpha-with-item-removed for item j is Cronbach's alpha (see below) recomputed using only the remaining items, so that a low value flags an item whose removal would increase the overall reliability of the test.
At the test level, two forms of Cronbach's alpha are always computed and reported. Raw alpha uses the standard variance-based formula
\alpha = \frac{k}{k - 1}\left(1 - \frac{\sum_{j} \sigma_j^2}{\sigma_X^2}\right),
where k is the number of items, \sigma_j^2 is the variance of
item j, and \sigma_X^2 is the variance of the total score X. Raw
alpha reflects the reliability of the actual (unweighted) total score
obtained by simply summing the item scores - the score most tests
actually use for reporting and decisions. Standardized alpha instead first
standardizes every item to unit variance before combining them,
\alpha_{std} = \frac{k \bar{r}}{1 + (k - 1)\bar{r}},
where \bar{r} is the average pairwise correlation among all items.
Standardized alpha differs meaningfully from raw alpha mainly when items
vary substantially in scale or format (e.g., a mix of dichotomous and
polytomous items with very different score ranges); when all items share
the same scale and format, the two are typically close, and raw alpha
remains the more directly interpretable of the two since it matches the
reliability of the score actually used in practice. See Cronbach (1951)
for the original derivation of coefficient alpha, and Osburn (2000) for a
discussion contrasting the raw (covariance-based) and standardized
(correlation-based) forms.
The standard error of measurement (SEM) is
SEM = SD(X)\sqrt{1 - \alpha} (using raw alpha), the standard CTT
relationship between test reliability and measurement precision.
The total-score frequency distribution is computed internally by
freq_score(), using the row sums of data after applying the same
missing recoding and listwise deletion of incomplete rows used
throughout the rest of this function, so that the frequency distribution
reflects exactly the same set of examinees and the same total-score
definition used elsewhere in the analysis. freq_score() is also
exported separately, for callers who only need a frequency distribution
for an arbitrary vector of integer total scores.
An object of class "ctt", a list with elements:
item |
A data frame with one row per item, containing the item
label, number of score categories, difficulty, the raw (uncorrected)
item-total correlation ( |
crit |
A list echoing the |
alpha |
A one-row test-level summary data frame containing
|
freq |
The total-score frequency distribution table returned by
|
call |
The matched call, as used by the |
Hwanggyu Lim hglim83@gmail.com
Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297-334. \Sexpr[results=rd]{tools:::Rd_expr_doi("10.1007/BF02310555")}.
Osburn, H. G. (2000). Coefficient alpha and related internal consistency reliability coefficients. Psychological Methods, 5(3), 343-355. \Sexpr[results=rd]{tools:::Rd_expr_doi("10.1037/1082-989X.5.3.343")}.
freq_score(), print.ctt(), summary.ctt()
# A small dichotomous example
set.seed(1)
dat <- data.frame(matrix(rbinom(300 * 8, 1, 0.6), nrow = 300))
out <- ctt(data = dat)
out
summary(out)
# A more realistic, mixed-format example: simulate response data for a
# 55-item test (50 dichotomous 3PLM items + 5 polytomous GRM items) from
# item parameters imported from a flexMIRT output file bundled with irtQ,
# then run ctt() on the simulated data
flex_sam <- system.file("extdata", "flexmirt_sample-prm.txt", package = "irtQ")
x <- bring.flexmirt(file = flex_sam, "par")$Group1$full_df
set.seed(2)
theta <- rnorm(1000)
dat_mixed <- simdat(x = x, theta = theta, D = 1)
out_mixed <- ctt(data = dat_mixed, item.id = x$id, cats = x$cats)
summary(out_mixed)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.