View source: R/percent_overlap.R
| estimate_overlap | R Documentation |
Unified front-end for KDE-based proportional overlap between two categories.
The returned overlap column is a 0–1 proportion, not a 0–100
percentage.
estimate_overlap(
data,
features,
category_col,
group_col = NULL,
min_tokens = 20,
bw = c("Hpi", "Hscv", "Hpi.diag", "scott.diag"),
eval_on = c("pooled", "group1", "group2", "pooled_sample"),
eval_n = NULL,
eval_seed = NULL,
engine = c("ks", "fast_diag", "fast_diagonal"),
chunk_size = 1000L,
method = c("mc", "legacy"),
density = c("kde", "mvnorm"),
mc_n = 10000L,
bw_scale = 1,
...
)
data |
Data frame with at least |
features |
Character vector of feature column names (e.g., c("F1","F2")). |
category_col |
Name of the column giving the two-way category factor. |
group_col |
Optional character vector of one or more grouping columns.
If provided, returns per-group overlap. Multiple grouping columns are
combined into a labeled |
min_tokens |
Minimum total tokens required (globally or per group). |
bw |
Bandwidth selection method passed to |
eval_on |
KDE evaluation points passed to |
eval_n |
Optional maximum number of KDE evaluation points. |
eval_seed |
Optional integer seed for KDE evaluation-point subsampling. |
engine |
KDE evaluation engine passed to |
chunk_size |
Chunk size for |
method |
Estimator passed to |
density |
Density model passed to |
mc_n |
Positive integer; number of Monte-Carlo samples drawn from each
fitted Gaussian when |
bw_scale |
Positive bandwidth multiplier passed to
|
... |
Additional arguments passed to |
A tibble (global = one row; grouped = one per group) with
overlap as a 0–1 proportion.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.