hierarchical_clust: Hierarchical Clustering on the Axes of an Analysis

View source: R/clust.R

hierarchical_clustR Documentation

Hierarchical Clustering on the Axes of an Analysis

Description

Clusters the individuals of a principal component analysis or of a multiple correspondence analysis, or the levels of one margin of a correspondence analysis, on the first 'ncp' axes. The clusters are those of FactoMineR::HCPC: Ward's hierarchical clustering, cut into 'nb_clust' clusters, then consolidated by k-means. Use it inside dplyr::mutate() to add them to the data frame:

'data <- data |> mutate(clust = hierarchical_clust(res, ncp = 3, nb_clust = 6))'

To choose the number of clusters, look at the tree first: 'hierarchical_clust(res, ncp = 3)'. The tree is built once: cutting it again, into another number of clusters or with names, is instant. Name the clusters with 'names', then describe them with clust_tab and draw them with ggfacto(clust = ).

Usage

hierarchical_clust(
  res,
  ncp,
  nb_clust = -1,
  tree = nb_clust == -1,
  consol = TRUE,
  margin = "rows",
  names = NULL
)

Arguments

res

An analysis made with multiple_correspondence_analysis, principal_component_analysis or correspondence_analysis (or with FactoMineR::MCA(), PCA() or CA(), or GDAtools::speMCA() or csMCA(), whose subcloud alone is clustered).

ncp

The number of axes to cluster on: the first ones, those worth interpreting (see the eigenvalues under interpret). There is no default: on every axis, the clustering would follow noise.

nb_clust

The number of clusters. With '-1', the default, the tree is cut where the gain in between-cluster inertia drops the most. Given 'names', it is the number of names.

tree

Should the clustering tree be drawn, to choose the number of clusters? By default, only when 'nb_clust = -1'. Its leaves are the distinct points of the cloud: the answer profiles of a multiple correspondence analysis, the levels of a correspondence analysis. Each bar of the inertia gains carries the number of clusters it makes, and the title the share of the inertia of the 'ncp' axes that lies between the clusters drawn.

consol

The k-means consolidation, which moves each individual to its nearest cluster once the tree is cut. 'TRUE', the default, is FactoMineR::HCPC's k-means, which counts every individual once, weights or not. '"weighted"' counts each by its weight (the survey weights, or a level's count in a correspondence analysis) with a simpler k-means, Lloyd's, which can place a few individuals differently even without weights. 'FALSE' keeps the clusters as the tree cuts them.

margin

For a correspondence analysis, the levels to cluster: '"rows"', the default, or '"columns"'.

names

The names of the clusters, in the order the levels should take: either 'c("Name 1", "Name 2", ...)', for clusters 1, 2, etc., or 'c("Name" = 2, "Other name" = 1, ...)', as in forcats::fct_recode(). Clusters are numbered along the first axis, so the names belong to one cut of one tree: write them after looking at that cut.

Details

The tree is kept in memory for the session, under a key made of everything it is built from — the coordinates on the 'ncp' axes, the weights, the answer profiles — so a tree made from other data is never reused. The last 20 trees are kept; 'options(ggfacto.clust_cache = 50)' keeps more, and 'options(ggfacto.clust_cache = 0)' none.

Value

A factor with the clusters, '"1"', '"2"', etc. or their 'names', numbered along the first axis and returned invisibly: a bare call only draws the tree. Inside dplyr::mutate(), it has one value per row of that data frame, and 'NA' on the rows the analysis did not use (when it was made on a subset of the population). Outside mutate(), it has one value per row of the data frame the analysis started from. For a correspondence analysis, each row of the data frame gets the cluster of its level ('NA' for a level outside the table); outside mutate(), there is one value per level, named after it.

Examples

data(tea, package = "FactoMineR")
res.mca <- multiple_correspondence_analysis(tea, 1:18)

# The tree, to choose the number of clusters, then the clusters, written into the data frame
hierarchical_clust(res.mca, ncp = 3)
tea <- tea |>
  dplyr::mutate(clust = hierarchical_clust(res.mca, ncp = 3, nb_clust = 6))

clust_tab(res.mca, tea, clust)
ggfacto(res.mca, tea, clust = clust)

# Named, in the order of your choice (the tree is not built again)
tea <- tea |>
  dplyr::mutate(clust = hierarchical_clust(res.mca, ncp = 3, names = c(
    "Cluster A" = 1, "Cluster B" = 2, "Cluster C" = 3, "Cluster E" = 5, "Cluster D" = 4,
    "Cluster F" = 6
  )))

# On a subset of the population: the other rows get NA
res.mca_young <- tea |>
  dplyr::filter(age < 30) |>
  multiple_correspondence_analysis(1:18)
tea <- tea |>
  dplyr::mutate(clust_young = hierarchical_clust(res.mca_young, ncp = 3, nb_clust = 4))

# A correspondence analysis clusters the levels of one margin: each individual gets the
# cluster of its level
res.ca <- forcats::gss_cat |>
  tabxplor::tab(relig, partyid) |>
  correspondence_analysis()
gss <- forcats::gss_cat |>
  dplyr::mutate(relig_clust = hierarchical_clust(res.ca, ncp = 2, nb_clust = 4))

ggfacto documentation built on Sept. 23, 2026, 1:08 a.m.