View source: R/HierarchicalClusterDists.R
| HierarchicalClusterDists | R Documentation |
Performs agglomerative hierarchical clustering from a precomputed dissimilarity object.
The function accepts either a dist object or a validated square
dissimilarity matrix. If ClusterNo > 0, the hierarchy is cut into the
requested number of clusters. If ClusterNo = 0, the complete
dendrogram is returned and is plotted only when PlotIt = TRUE.
This is primarily an internal helper. For the usual user interface, see
HierarchicalClustering.
HierarchicalClusterDists(
pDist,
ClusterNo = 0,
Type = "ward.D2",
ColorTreshold = 0,
Fast = FALSE,
PlotIt=FALSE,
...
)
pDist |
Precomputed pairwise dissimilarities supplied either as an object of class
A raw feature matrix is not converted to distances by this function.
Distances must be calculated beforehand, for example with
For |
ClusterNo |
Integer between zero and the number of observations. If If |
Type |
Character scalar specifying the agglomeration method. Supported values are
See |
ColorTreshold |
Finite numeric scalar.
If nonzero and If the fusion heights are non-monotone, only a horizontal reference line
is drawn and Height scales depend on |
Fast |
Logical scalar. If |
PlotIt |
If |
... |
Further arguments passed to a selected specialized method.
In the ordinary |
Input validation
If pDist is supplied as a matrix or data frame, it is converted to a
matrix and checked to be numeric, square, at least 2 \times 2, finite,
symmetric, non-negative within floating-point tolerance, and to have a zero
diagonal. Negligible asymmetry and small negative round-off values are
corrected before conversion to a dist object.
If a dist object is supplied directly, its size, expected vector
length, finiteness, and non-negativity are validated.
Ward variants
The Ward methods have distinct interpretations:
"ward.D" applies R's legacy Ward Lance-Williams update directly
to the supplied dissimilarities. It should not be interpreted as the
generalized Ward pseudo-inertia criterion for arbitrary Minkowski
dissimilarities.
"ward.D2" uses R's Ward.D2 implementation. The supplied
dissimilarities are squared internally before the Ward update. For arbitrary
dissimilarities, this can be interpreted as a Ward-like hierarchy based on
squared dissimilarities. Reported heights are on a square-root,
distance-like scale.
"ward.pseudo" implements the dissimilarity-based generalized
Ward criterion described by Chavent et al. (2018).
For "ward.pseudo", let w_i denote the observation weights and
\mu_C = \sum_{i \in C} w_i the weight of cluster C. The
dissimilarity-based pseudo-inertia is
I_D(C) =
\frac{
\sum_{i,j \in C} w_i w_j d_{ij}^2
}{
2 \mu_C
}.
At every agglomeration step, clusters A and B minimizing
\Delta_D(A,B) =
I_D(A \cup B) - I_D(A) - I_D(B)
are merged.
This implementation uses equal normalized weights
w_i = \frac{1}{n}.
Therefore the singleton aggregation costs are
\Delta_{ij} = \frac{d_{ij}^2}{2n}.
These transformed aggregation costs, not the original dissimilarities, are
passed internally to the "ward.D" Lance-Williams update together with
members = rep(1/n, n).
The generalized pseudo-inertia criterion is defined for potentially non-Euclidean dissimilarities.
For equal observation weights, "ward.pseudo" and
"ward.D2" applied to the same original dissimilarities produce the same
sequence of cluster mergers, apart from possible ordering differences caused
by exact numerical ties. However, their height scales differ according to
h_{\mathrm{ward.pseudo}} =
\frac{h_{\mathrm{ward.D2}}^2}{2n}.
This rescaling can make an appropriate dendrogram cut easier to identify visually.
The "ward.pseudo" option is not the Ward_p algorithm of de Amorim
(2015), which additionally uses cluster-dependent feature weights and
L_p cluster centers.
Plotting and non-monotone heights
When ClusterNo = 0 and PlotIt = TRUE, the dendrogram is plotted. For
"ward.pseudo", the y-axis represents the increase in pseudo
within-cluster inertia; for other methods it is the ordinary fusion height.
If ColorTreshold != 0 and fusion heights are monotone,
rect.hclust() is used. If they are non-monotone, a horizontal red line
is drawn instead and a warning recommends using ClusterNo rather than a
height cut to obtain a partition.
A list with components:
Cls |
If |
Dendrogram |
Object of class |
Object |
Object of class For |
A raw n \times d feature matrix must not be passed as pDist.
Compute a dissimilarity object beforehand.
For example, a Minkowski p = 6 hierarchy using the generalized
pseudo-inertia criterion can be obtained from
stats::dist(X, method = "minkowski", p = 6) followed by
Type = "ward.pseudo".
Michael Thrun
Chavent, M., Kuentz-Simonet, V., Labenne, A., and Saracco, J. (2018). ClustGeo: an R package for hierarchical clustering with spatial constraints. Computational Statistics, 33, 1799–1822. doi: 10.1007/s00180-018-0791-1.
Murtagh, F. and Legendre, P. (2014). Ward's hierarchical agglomerative clustering method: Which algorithms implement Ward's criterion? Journal of Classification, 31, 274–295. doi: 10.1007/s00357-014-9161-z.
de Amorim, R. C. (2015). Feature relevance in Ward's hierarchical clustering using the Lp norm. Journal of Classification, 32, 46–62. doi: 10.1007/s00357-015-9167-1.
HierarchicalClusterData,
HierarchicalClustering,
ClusterDendrogram,
hclust,
cutree
## Not run:
data("Hepta")
# Standard Ward.D2 clustering
D <- stats::dist(Hepta$Data)
res <- HierarchicalClusterDists(
D,
ClusterNo = 7,
Type = "ward.D2"
)
res$Cls
# Minkowski p = 6 with generalized Ward pseudo-inertia
Xscaled <- scale(Hepta$Data)
D6 <- stats::dist(
Xscaled,
method = "minkowski",
p = 6
)
res_pseudo <- HierarchicalClusterDists(
D6,
ClusterNo = 6,
Type = "ward.pseudo"
)
# Plot the complete hierarchy
HierarchicalClusterDists(
D,
ClusterNo = 0,
Type = "complete"
)
## End(Not run)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.