R/data.R

#' IPIP Big Five 50-Item Inventory with Qwen3 Embeddings
#'
#' A bundled example dataset containing the 50-item IPIP Big Five personality
#' inventory with precomputed sentence-embedding vectors. The scale has 5 factors
#' (Extraversion, Agreeableness, Conscientiousness, Neuroticism, Openness) with
#' 10 items each, including 18 reverse-keyed items, making it suitable for
#' demonstrating all encoding methods.
#'
#' @format A list with components:
#' \describe{
#'   \item{items}{Character vector (length 50): item text.}
#'   \item{codes}{Character vector (length 50): item codes (E1, E2, ..., O50).}
#'   \item{factors}{Character vector (length 50): theoretical factor labels.}
#'   \item{scoring}{Numeric vector (length 50): +1 or -1 keying direction.}
#'   \item{embeddings}{Numeric matrix (50 x 4096): precomputed embeddings from
#'     the \code{Qwen3-Embedding-8B} model, rounded to 4 decimal places.}
#' }
#'
#' @source Items from the International Personality Item Pool (IPIP;
#'   \url{https://ipip.ori.org/}), which is in the public domain. Embeddings were
#'   generated with the \code{Qwen/Qwen3-Embedding-8B} model
#'   (\url{https://huggingface.co/Qwen/Qwen3-Embedding-8B})
#'   and rounded to 4 decimal places to reduce file size. The regeneration
#'   script is in \code{data-raw/big5.R}.
#'
#' @examples
#' data(big5)
#' str(big5)
#' table(big5$factors, big5$scoring)
"big5"

Try the semanticfa package in your browser

Any scripts or data that you put into this service are public.

semanticfa documentation built on Sept. 2, 2026, 1:07 a.m.