HierarchicalClusterDists: Internal Hierarchical Clustering from Precomputed...

View source: R/HierarchicalClusterDists.R

HierarchicalClusterDistsR Documentation

Internal Hierarchical Clustering from Precomputed Dissimilarities

Description

Performs agglomerative hierarchical clustering from a precomputed dissimilarity object.

The function accepts either a dist object or a validated square dissimilarity matrix. If ClusterNo > 0, the hierarchy is cut into the requested number of clusters. If ClusterNo = 0, the complete dendrogram is returned and is plotted only when PlotIt = TRUE.

This is primarily an internal helper. For the usual user interface, see HierarchicalClustering.

Usage

HierarchicalClusterDists(
  pDist,
  ClusterNo = 0,
  Type = "ward.D2",
  ColorTreshold = 0,
  Fast = FALSE,
  PlotIt=FALSE,
  ...
)

Arguments

pDist

Precomputed pairwise dissimilarities supplied either as an object of class "dist" or as a finite, non-negative, symmetric numeric n \times n matrix with zero diagonal.

A raw feature matrix is not converted to distances by this function. Distances must be calculated beforehand, for example with stats::dist().

For Type = "ward.pseudo", pDist must contain the original unsquared dissimilarities.

ClusterNo

Integer between zero and the number of observations.

If ClusterNo = 0, the complete hierarchy is returned. It is plotted only when PlotIt = TRUE; Cls = NULL is returned unless a height threshold is used.

If ClusterNo > 0, the hierarchy is cut into exactly ClusterNo clusters using stats::cutree().

Type

Character scalar specifying the agglomeration method. Supported values are "ward.D", "ward.D2", "ward.pseudo", "single", "complete", "average", "mcquitty", "median", "centroid", "MinEnergy", "Minimax", "Sparse", "Gini", and its synonym "Genie".

See Details for the distinction between the three Ward variants.

ColorTreshold

Finite numeric scalar. If nonzero and ClusterNo = 0, the value is interpreted on the dendrogram height scale. For a monotone ultramteric portion of given dissimilarity, stats::rect.hclust() draws rectangles at this height and classifiation vector Cls is return based on classification vector Cls is returned from the corresponding cut.

If the fusion heights are non-monotone, only a horizontal reference line is drawn and Cls = NULL.

Height scales depend on Type; in particular, a threshold used for "ward.D2" cannot be transferred unchanged to "ward.pseudo".

Fast

Logical scalar. If TRUE and package fastcluster is installed, fastcluster::hclust() is used. Otherwise stats::hclust() is used. If Fast = TRUE but fastcluster is unavailable, a warning is issued.

PlotIt

If TRUE and ClusterNo > 0, plots the dendrogram using ClusterDendrogram. If TRUE and ClusterNo = 0, plots the complete dendrogram using base R graphics. The default is FALSE.

...

Further arguments passed to a selected specialized method. In the ordinary hclust path, these arguments are currently not used.

Details

Input validation

If pDist is supplied as a matrix or data frame, it is converted to a matrix and checked to be numeric, square, at least 2 \times 2, finite, symmetric, non-negative within floating-point tolerance, and to have a zero diagonal. Negligible asymmetry and small negative round-off values are corrected before conversion to a dist object.

If a dist object is supplied directly, its size, expected vector length, finiteness, and non-negativity are validated.

Ward variants

The Ward methods have distinct interpretations:

  • "ward.D" applies R's legacy Ward Lance-Williams update directly to the supplied dissimilarities. It should not be interpreted as the generalized Ward pseudo-inertia criterion for arbitrary Minkowski dissimilarities.

  • "ward.D2" uses R's Ward.D2 implementation. The supplied dissimilarities are squared internally before the Ward update. For arbitrary dissimilarities, this can be interpreted as a Ward-like hierarchy based on squared dissimilarities. Reported heights are on a square-root, distance-like scale.

  • "ward.pseudo" implements the dissimilarity-based generalized Ward criterion described by Chavent et al. (2018).

For "ward.pseudo", let w_i denote the observation weights and \mu_C = \sum_{i \in C} w_i the weight of cluster C. The dissimilarity-based pseudo-inertia is

I_D(C) = \frac{ \sum_{i,j \in C} w_i w_j d_{ij}^2 }{ 2 \mu_C }.

At every agglomeration step, clusters A and B minimizing

\Delta_D(A,B) = I_D(A \cup B) - I_D(A) - I_D(B)

are merged.

This implementation uses equal normalized weights

w_i = \frac{1}{n}.

Therefore the singleton aggregation costs are

\Delta_{ij} = \frac{d_{ij}^2}{2n}.

These transformed aggregation costs, not the original dissimilarities, are passed internally to the "ward.D" Lance-Williams update together with members = rep(1/n, n).

The generalized pseudo-inertia criterion is defined for potentially non-Euclidean dissimilarities.

For equal observation weights, "ward.pseudo" and "ward.D2" applied to the same original dissimilarities produce the same sequence of cluster mergers, apart from possible ordering differences caused by exact numerical ties. However, their height scales differ according to

h_{\mathrm{ward.pseudo}} = \frac{h_{\mathrm{ward.D2}}^2}{2n}.

This rescaling can make an appropriate dendrogram cut easier to identify visually.

The "ward.pseudo" option is not the Ward_p algorithm of de Amorim (2015), which additionally uses cluster-dependent feature weights and L_p cluster centers.

Plotting and non-monotone heights

When ClusterNo = 0 and PlotIt = TRUE, the dendrogram is plotted. For "ward.pseudo", the y-axis represents the increase in pseudo within-cluster inertia; for other methods it is the ordinary fusion height.

If ColorTreshold != 0 and fusion heights are monotone, rect.hclust() is used. If they are non-monotone, a horizontal red line is drawn instead and a warning recommends using ClusterNo rather than a height cut to obtain a partition.

Value

A list with components:

Cls

If ClusterNo > 0, a numerical vector of length n defining the classification. If ClusterNo = 0, NULL unless a height threshold is used.

Dendrogram

Object of class "dendrogram" representing the complete hierarchy.

Object

Object of class "hclust".

For Type = "ward.pseudo", Object$height contains the successive increases in normalized pseudo within-cluster inertia.

Note

A raw n \times d feature matrix must not be passed as pDist. Compute a dissimilarity object beforehand.

For example, a Minkowski p = 6 hierarchy using the generalized pseudo-inertia criterion can be obtained from stats::dist(X, method = "minkowski", p = 6) followed by Type = "ward.pseudo".

Author(s)

Michael Thrun

References

Chavent, M., Kuentz-Simonet, V., Labenne, A., and Saracco, J. (2018). ClustGeo: an R package for hierarchical clustering with spatial constraints. Computational Statistics, 33, 1799–1822. doi: 10.1007/s00180-018-0791-1.

Murtagh, F. and Legendre, P. (2014). Ward's hierarchical agglomerative clustering method: Which algorithms implement Ward's criterion? Journal of Classification, 31, 274–295. doi: 10.1007/s00357-014-9161-z.

de Amorim, R. C. (2015). Feature relevance in Ward's hierarchical clustering using the Lp norm. Journal of Classification, 32, 46–62. doi: 10.1007/s00357-015-9167-1.

See Also

HierarchicalClusterData, HierarchicalClustering, ClusterDendrogram, hclust, cutree

Examples

## Not run: 
data("Hepta")

# Standard Ward.D2 clustering
D <- stats::dist(Hepta$Data)

res <- HierarchicalClusterDists(
  D,
  ClusterNo = 7,
  Type = "ward.D2"
)

res$Cls

# Minkowski p = 6 with generalized Ward pseudo-inertia
Xscaled <- scale(Hepta$Data)

D6 <- stats::dist(
  Xscaled,
  method = "minkowski",
  p = 6
)

res_pseudo <- HierarchicalClusterDists(
  D6,
  ClusterNo = 6,
  Type = "ward.pseudo"
)

# Plot the complete hierarchy
HierarchicalClusterDists(
  D,
  ClusterNo = 0,
  Type = "complete"
)

## End(Not run)

FCPS documentation built on Oct. 3, 2026, 9:06 a.m.