tabscale: Comprehensive Scale Analysis

tabscaleR Documentation

Comprehensive Scale Analysis

Description

tabscale() provides a one-command, publication-ready psychometric report. It covers internal consistency, stability, equivalence, inter-rater reliability, measurement error, content validity, structural validity, convergent/discriminant and known-groups validity, criterion validity, measurement invariance, DIF screening, and responsiveness when the required data are supplied. Variable labels are used throughout whenever available.

Usage

tabscale(data = NULL, vars = NULL, factor = NULL, reverse = NULL,
         report = c("auto", "brief", "full", "custom"),
         range = NULL, score = c("mean", "sum"), min_valid = NULL,
         missing = c("pairwise", "complete"), cor_method = c("pearson", "spearman"),
         reliability = TRUE, bootstrap = 0, conf = 0.95,
         retest = NULL, parallel_form = NULL, raters = NULL,
         rater_type = c("auto", "continuous", "categorical"),
         content = NULL, content_cutoff = 3, content_max = 4,
         validity = TRUE,
         efa = NULL, efa_method = c("pa", "ml"), nfactor = NULL,
         rotation = c("varimax", "promax", "none"), parallel_iter = 100,
         cfa = NULL, ordered = FALSE, estimator = "auto",
         cfa_missing = "fiml", cfa_group = NULL,
         modification = FALSE, modification_min = 10,
         invariance = FALSE,
         invariance_levels = c("configural", "metric", "scalar", "strict"),
         dif = FALSE, convergent = NULL, discriminant = NULL,
         convergent_min = 0.50, discriminant_max = 0.30,
         known_groups = NULL,
         gold = NULL, event = NULL, direction = c("auto", "higher", "lower"),
         post = NULL,
         name = FALSE, digit = 2, p_digit = 3,
         template = c("journal", "clean", "minimal"),
         plot = TRUE, plot_types = "auto",
         viewer_plot_format = c("png", "svg"),
         append = NULL, file = NULL,
         title = NULL, interpretation = FALSE,
         raw = TRUE, show = TRUE, seed = NULL)

Arguments

data

Optional data frame. When omitted, the active R4VN data frame is used.

vars

Items created by vars(), a character vector, or a one-sided formula.

factor

Optional named list defining subscales/CFA factors.

reverse

Optional items to reverse-score.

report

Output profile. "auto" runs reliability plus EFA and any data-dependent modules requested by their arguments; "brief" omits factor models; "full" also runs CFA when a factor map and lavaan are available; "custom" follows the module switches exactly.

range

Two numeric values giving the minimum and maximum item score.

score

Calculate scale scores as the item "mean" or "sum".

min_valid

Minimum valid items. A value in ⁠(0,1]⁠ is treated as a proportion. A named vector can define separate minima for factors and Total.

missing

Correlation/covariance handling: pairwise or complete observations.

cor_method

Pearson or Spearman item correlations.

reliability

Logical; calculate reliability statistics.

bootstrap

Number of nonparametric bootstrap replicates for alpha and omega-total confidence intervals. Zero uses a Feldt interval for alpha.

conf

Confidence level for alpha and AUC intervals.

retest

Items measured again, in the same order as vars(), for test-retest reliability, ICC, SEM, MDC, and Bland-Altman analysis.

parallel_form

Items from an equivalent form, in the same order as vars(), for parallel-form reliability.

raters

Two or more variables containing ratings of the same subjects.

rater_type

Treat ratings as continuous or categorical; "auto" chooses categorical for variables with at most 10 observed levels.

content

Expert-by-item matrix/data frame of content-relevance ratings.

content_cutoff

Minimum rating counted as content-relevant.

content_max

Maximum possible content rating, retained in the report.

validity

Logical master switch for validity modules.

efa

Logical or NULL; run EFA, KMO, Bartlett, and parallel analysis.

efa_method

Principal-axis ("pa") or maximum-likelihood ("ml") EFA.

nfactor

Number of EFA factors. When NULL, parallel analysis is used.

rotation

EFA rotation.

parallel_iter

Number of Monte Carlo samples for parallel analysis.

cfa

Logical or NULL; run CFA using the optional lavaan package.

ordered

Logical or character item names treated as ordinal in CFA.

estimator

CFA estimator. "auto" uses WLSMV for ordered items and MLR otherwise.

cfa_missing

Missing-data option passed to lavaan for non-ordinal CFA.

cfa_group

Optional grouping variable name for multiple-group CFA.

modification

Logical; include large CFA modification indices.

modification_min

Minimum modification index displayed.

invariance

Logical; test configural, metric, scalar, and strict measurement invariance across cfa_group.

invariance_levels

Invariance levels to fit.

dif

Logical; screen uniform and non-uniform differential item functioning across a two-level cfa_group.

convergent

External variables used for convergent validity.

discriminant

External variables used for discriminant validity.

convergent_min

Prespecified minimum absolute convergent correlation.

discriminant_max

Prespecified maximum absolute discriminant correlation.

known_groups

Grouping variable for known-groups validity.

gold

Optional criterion or binary gold-standard variable.

event

Event level for binary gold-standard ROC analysis.

direction

Whether higher or lower scores predict the event; "auto" chooses the direction with AUC at least 0.5.

post

Post-intervention/follow-up items, in the same order as vars(), for responsiveness (effect size and standardized response mean).

name

FALSE does not modify data. TRUE creates scale_total and subscale variables; a character value supplies the score prefix.

digit

Decimal places for estimates.

p_digit

Decimal places for p-values.

template

HTML style.

plot

Logical; create all applicable graphics in both the R Plots pane and the HTML Viewer.

plot_types

"auto", "all", or any of "items", "distributions", "correlation", "reliability", "scores", "scree", "loadings", "cfa_loadings", "roc", "test_retest", "parallel_form", "inter_rater", "content", "external_validity", "known_groups", "invariance", and "responsiveness". The automatic set omits secondary plots that often add visual clutter: item completeness, item-response distributions, item-reliability diagnostics, external-validity correlations, known-groups boxplots, and individual responsiveness trajectories. Their statistical tables remain in the report. Request any of these names explicitly, or use plot_types = "all", when the graphic is needed.

viewer_plot_format

Format used to embed plots in the HTML Viewer. The default "png" uses high-resolution, self-contained images and gives consistent fonts in RStudio Viewer, Chrome, and saved HTML files. "svg" retains vector graphics but may render text differently across browsers because SVG font substitution is controlled by the local system.

append

Optional previous R4VN table object or HTML file.

file

Optional HTML output path.

title

Optional table title.

interpretation

Logical; add cautious automatic interpretation. The default is FALSE.

raw

Logical; retain numerical result components.

show

Logical; open the HTML result.

seed

Optional random seed used by parallel analysis. The default NULL does not set a seed.

Details

The default report = "auto" produces descriptive item distributions, missing/floor/ceiling effects, corrected item-total correlations, alpha with 95% CI, standardized alpha, omega total, split-half coefficients, all six Guttman lambdas, KMO, Bartlett's test, parallel analysis, EFA, score distributions, and matching graphics. KR-20 is added for binary items. Ordinal alpha, omega hierarchical, and the greatest lower bound are added when the optional package psych is installed.

Additional data activate stability/test-retest, parallel-form, inter-rater, content, convergent, discriminant, known-groups, criterion, responsiveness, measurement-invariance, and DIF sections. CFA and invariance use lavaan. Face validity is inherently qualitative and is therefore identified in the coverage table rather than assigned a spurious numeric coefficient.

The same plot specifications are rendered in the R graphics device and in the HTML Viewer. In RStudio, use the Plots pane arrows to review every graph, or rerun selected graphs with plot(result, which = "roc").

Value

Invisibly returns an object inheriting from r4vn_tabscale and r4vn_tab. Components include tables, coverage, descriptive, reliability, scores, test_retest, parallel_form, inter_rater, content_validity, efa, cfa, convergent_validity, discriminant_validity, known_groups, criterion_validity, invariance, dif, responsiveness, plots, interpretation, and file.

See Also

vars, tab, tabmulti, tabexport

Other R4VN tables: tab(), tabexport(), tabforest(), tablong(), tabmeta(), tabmulti(), tabscore(), tabsurvey(), vars()

Examples

set.seed(2026)
n <- 120
f1 <- rnorm(n)
f2 <- 0.35 * f1 + rnorm(n, sd = 0.94)
make_item <- function(z) as.integer(cut(z, quantile(z, 0:5/5),
                                       include.lowest = TRUE, labels = FALSE))
latent <- list(0.8*f1, 0.7*f1, 0.9*f1, -0.7*f1,
               0.8*f2, 0.7*f2, 0.9*f2, 0.6*f2)
base_items <- lapply(latent, function(z) make_item(z + rnorm(n)))
dat <- as.data.frame(base_items)
names(dat) <- paste0("q", 1:8)
for (j in 1:8) {
  dat[[paste0("q", j, "_retest")]] <- pmax(
    1,
    pmin(
      5,
      dat[[paste0("q", j)]] + sample(-1:1, n, TRUE, c(.1, .8, .1))
    )
  )
  dat[[paste0("q", j, "_formb")]] <- make_item(latent[[j]] + rnorm(n))
  dat[[paste0("q", j, "_post")]] <- pmax(1, pmin(5, dat[[paste0("q",j)]] + rbinom(n,1,.35)))
  attr(dat[[paste0("q",j)]], "label") <- paste("Well-being item", j)
}
dat$convergent_measure <- f1 + f2 + rnorm(n, sd=.6)
dat$unrelated_measure <- rnorm(n)
dat$known_group <- factor(ifelse(f1+f2>0,"Higher expected score","Lower expected score"))
dat$gold <- factor(ifelse(f1 + f2 + rnorm(n) > 0, "Yes", "No"),
                   levels = c("No", "Yes"))
dat$rater1 <- sample(1:4,n,TRUE); dat$rater2 <- dat$rater1
dat$rater3 <- dat$rater1
dat$rater2[sample(n,30)] <- sample(1:4,30,TRUE)
dat$rater3[sample(n,35)] <- sample(1:4,35,TRUE)
attr(dat$known_group,"label") <- "Prespecified clinical group"
attr(dat$gold,"label") <- "Clinical gold standard"

# 1. One-command automatic report: reliability, factorability, EFA, plots.
tb <- tabscale(
  dat,
  vars = vars(q1, q2, q3, q4, q5, q6, q7, q8),
  factor = list(Domain1 = vars(q1, q2, q3, q4),
                Domain2 = vars(q5, q6, q7, q8)),
  reverse = vars(q4), range = c(1, 5),
  nfactor = 2, parallel_iter = 10,
  plot = FALSE, show = FALSE
)
tb$reliability_summary
tb$efa$loadings
if (interactive()) {
  plot(tb)                         # all plots in the Plots pane
  plot(tb, which = "reliability") # one selected plot
}

# 2. Stability, parallel forms, inter-rater reliability, and measurement error.
rel <- tabscale(
  dat, vars=vars(q1,q2,q3,q4,q5,q6,q7,q8), reverse=vars(q4), range=c(1,5),
  retest=vars(q1_retest,q2_retest,q3_retest,q4_retest,
              q5_retest,q6_retest,q7_retest,q8_retest),
  parallel_form=vars(q1_formb,q2_formb,q3_formb,q4_formb,
                     q5_formb,q6_formb,q7_formb,q8_formb),
  raters=vars(rater1,rater2,rater3),
  plot=FALSE, show=FALSE)
rel$test_retest$table
rel$parallel_form$table
rel$inter_rater$table

# 3. Convergent, discriminant, known-groups, criterion validity,
#    responsiveness, and automatic ROC curves.
val <- tabscale(
  dat, vars=vars(q1,q2,q3,q4,q5,q6,q7,q8), reverse=vars(q4), range=c(1,5),
  convergent=vars(convergent_measure), discriminant=vars(unrelated_measure),
  known_groups=known_group, gold=gold, event="Yes",
  post=vars(q1_post,q2_post,q3_post,q4_post,q5_post,q6_post,q7_post,q8_post),
  interpretation=TRUE, plot=FALSE, show=FALSE)
val$convergent_validity
val$discriminant_validity
val$known_groups$tests
val$criterion_validity$table
val$responsiveness$table

# 4. Content validity: experts in rows and items in columns.
expert_ratings <- as.data.frame(matrix(sample(2:4, 6*8, TRUE,
  prob=c(.10,.30,.60)), nrow=6, dimnames=list(NULL,paste0("q",1:8))))
content_result <- tabscale(
  dat, vars=vars(q1,q2,q3,q4,q5,q6,q7,q8), range=c(1,5),
  content=expert_ratings, content_cutoff=3,
  plot=FALSE, show=FALSE)
content_result$content_validity$summary
content_result$content_validity$item

# 5. CFA, composite reliability, AVE, HTMT, Fornell-Larcker,
#    measurement invariance, and DIF screening.
if (interactive() && requireNamespace("lavaan", quietly = TRUE)) {
  cfa_result <- tabscale(
    dat, vars=vars(q1,q2,q3,q4,q5,q6,q7,q8),
    factor=list(Domain1=vars(q1,q2,q3,q4), Domain2=vars(q5,q6,q7,q8)),
    reverse=vars(q4), range=c(1,5), cfa=TRUE, ordered=TRUE,
    cfa_group=known_group, invariance=TRUE, dif=TRUE,
    modification=TRUE, plot=FALSE, show=FALSE)
  cfa_result$cfa$reliability
  cfa_result$cfa$htmt
  cfa_result$invariance$table
  cfa_result$dif
}

R4VN documentation built on Sept. 30, 2026, 5:13 p.m.