clust_tab: Describe Clusters with One Table

View source: R/clust.R

clust_tabR Documentation

Describe Clusters with One Table

Description

One table describing every cluster: each variable's levels down the page, the clusters across it, and a colour saying at a glance which levels a cluster is made of. Give it the analysis, the data frame and the clusters, in the order of ggmca: the active variables, the weights and the rows the analysis was made on are its own. Numeric variables come in as mean rows, coloured by their difference to the mean in standard deviations; the last two rows give each cluster's share of the population and its size (under the table when every row is a mean, as in a principal component analysis).

Usage

clust_tab(
  res,
  data,
  clust,
  row_vars,
  pct = "col",
  excl = NA,
  color = "difference",
  row_tot = "% of population",
  cleannames = TRUE,
  ...,
  wt
)

Arguments

res

The analysis the clusters were made on, with multiple_correspondence_analysis or principal_component_analysis (or FactoMineR::MCA() or PCA(), or GDAtools::speMCA() or csMCA()). For a correspondence analysis, cross the clusters with the other variable of the table with tabxplor::tab() instead.

data

The data frame, with the clusters. The whole data frame will do when the analysis was made on a subset of it: only the rows the analysis used are described.

clust

The variable with the clusters, typically made with hierarchical_clust, as a bare name or a string.

row_vars

<tidy-select> The variables to describe the clusters with: by default, the active variables of the analysis. Numeric ones become mean rows, unless 'shape' (passed on to tabxplor::tab()) cuts them into levels, e.g. 'shape = "sd_bands"', or 'shape = c(AGE = "quintiles")' for that one.

pct

'"col"' (default) reads each cluster as a distribution: of the people in this cluster, what percentage are in this level. '"row"' reads each level as a distribution across clusters.

excl

The levels not to show, matched exactly by name; their individuals still count in the percentages. 'NA', the default, hides the missing values (and the levels named '<VAR>.NA'); 'excl = NULL' shows every level.

color

The colour measure, see tab. With '"difference"' (the default), percentages are coloured by their difference with the whole population, and means by their difference in standard deviations — but not both in one table, which a single ladder cannot grade: there, means stay uncoloured, and '"ratio"' colours every row.

row_tot

The name of the row giving each cluster's share of the population.

cleannames

Set to FALSE to keep the level and cluster names as they are, prefix numbers like "1-" and text in parentheses included.

...

Additional arguments to pass to tab.

wt

Not used: the table is weighted with the weights of the analysis. It is there for the former form, 'clust_tab(data, row_vars, clust, wt)', still read as HCPC_tab.

Value

A tabxplor table — see [ggfacto_summary] for how it prints.

See Also

[ggfacto_summary], [hierarchical_clust()], [interpret()].

Examples

data(tea, package = "FactoMineR")
res.mca <- multiple_correspondence_analysis(tea, 1:18)
tea <- tea |>
  dplyr::mutate(clust = hierarchical_clust(res.mca, ncp = 3, nb_clust = 6))

# ONE option decides how every tabxplor table prints, an interpretation table included.
# In a script it goes once, at the top, beside the library() calls.
options(tabxplor.print = "html")

# The clusters, by the active variables
clust_tab(res.mca, tea, clust)

# ... and by other variables
clust_tab(res.mca, tea, clust, row_vars = c(sex, SPC, age), pct = "row")

# A principal component analysis: the means of each cluster
res.pca <- principal_component_analysis(mtcars, 1:7)
cars <- mtcars |>
  dplyr::mutate(clust = hierarchical_clust(res.pca, ncp = 2, nb_clust = 3))
clust_tab(res.pca, cars, clust)

ggfacto documentation built on Sept. 23, 2026, 1:08 a.m.