| multiple_correspondence_analysis | R Documentation |
A user-friendly wrapper around MCA, made to
work with ggfacto functions like ggmca, interpret and
hierarchical_clust. Variables are selected the way of the 'tidyverse', as in
tabxplor::tab(). Supplementary variables are not given here: they are added afterwards,
in ggmca.
'MCA2()' keeps the fit of ggfacto 0.3.2, on the individuals: '$ind' has one row per analysed row, so that 'FactoMineR::HCPC()' of it classifies the individuals, in their order. It gives the same graphs, tables and clusters as 'multiple_correspondence_analysis()', more slowly on large data, and will be deprecated.
multiple_correspondence_analysis(
data,
active_vars,
wt,
excl = NA,
ncp = Inf,
graph = FALSE,
filter,
...
)
MCA2(data, active_vars, wt, excl = NA, ncp = Inf, graph = FALSE, filter, ...)
data |
The data frame. To analyse a subset of the population, give the whole data frame and
'filter', or filter it inside the call with the native pipe,
'data |> dplyr::filter(...) |> multiple_correspondence_analysis(...)': the analysis then
remembers which rows it used, so that |
active_vars |
<tidy-select> The active variables. |
wt |
<tidy-select> The weight variable, if any. |
excl |
The levels to exclude from the calculation of the axes (specific multiple correspondence analysis), matched exactly by name. The missing values of each active variable become a level named '<VAR>.NA', and 'NA', the default, excludes all of them: 'excl = NA' for missing values only, 'excl = c(NA, "Other")' to exclude a level too, 'excl = "DIPLOMA.NA"' for the missing values of one variable only, 'excl = NULL' to keep every level. |
ncp |
The number of axes to keep. All of them by default: the eigenvalue table is how one
chooses how many axes to interpret, and a truncated one cannot show the drop — it also
renormalises Benzecri's modified rate over the axes it kept, so the same axis gets a different
rate. To cluster on the first axes, give |
graph |
By default no graph is made, since the result can be plotted with
|
filter |
A condition on the rows of 'data', as in |
... |
Additional arguments to pass to |
A 'MCA' object from FactoMineR, fitted on the distinct answer profiles (the
combinations of active answers), each weighted by its individuals: the eigenvalues and every
result on the levels are the individuals', and '$ind' has one row per profile (per individual
with 'MCA2()'). Use axis_coord and hierarchical_clust to write
coordinates and clusters into the data frame ('FactoMineR::HCPC()' would cluster the profiles).
One more element, 'source', records for each row of 'data' its row of '$ind' ('NA' if it was not
analysed) and its weight.
data(tea, package = "FactoMineR")
res.mca <- multiple_correspondence_analysis(tea, 1:18)
interpret(res.mca) # the eigenvalues, then the axes
ggfacto(res.mca, tea, sup_vars = c(sex, SPC)) # the graph, with supplementary variables
ggfacto(res.mca, tea, sup_vars = c(sex, SPC), interactive = TRUE) # hover: the crosstables
# A subset of the population: the analysis remembers which rows it used
res.mca_young <- tea |>
dplyr::filter(age < 30) |>
multiple_correspondence_analysis(1:18)
# the same analysis
res.mca_young <- multiple_correspondence_analysis(tea, 1:18, filter = age < 30)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.