HDDClustering: HDD clustering is a model-based clustering method of...

View source: R/HDDClustering.R

HDDClusteringR Documentation

HDD clustering is a model-based clustering method of [Bouveyron et al., 2007].

Description

HDD clustering is based on the Gaussian Mixture Model and on the idea that the data lives in subspaces with a lower dimension than the dimension of the original space. It uses the EM algorithm to estimate the parameters of the model [Berge et al., 2012].

Usage

HDDClustering(Data, ClusterNo, PlotIt=F,...)

Arguments

Data

[1:n,1:d] matrix containing the dataset to be clustered. It consists of n cases of d-dimensional data points. Every case has d attributes, variables or features.

ClusterNo

Optional numeric value indicating the number of clusters, or a vector from 1 to k indicating the maximum expected number of clusters.

PlotIt

(optional) Boolean. Default = FALSE = No plotting performed.

...

Further arguments to be set for the clustering algorithm, if not set, default arguments are used, see hddc for details.

Details

HDD clustering maximizes the BIC criterion over possible numbers of clusters up to ClusterNo. By default, the most general model is used. Alternatively, model = "ALL" evaluates all available models with BIC [Berge et al., 2012]. If specific properties of Data are known beforehand, see hddc for model selection.

Value

List of

Cls

[1:n] numerical vector with n numbers defining the classification as the main output of the clustering algorithm. It has k unique numbers representing the arbitrary labels of the clustering.

Object

Object defined by clustering algorithm as the other output of this algorithm

Author(s)

Quirin Stier

References

[Berge et al., 2012] L. Berge, C. Bouveyron and S. Girard, HDclassif: an R Package for Model-Based Clustering and Discriminant Analysis of High-Dimensional Data, Journal of Statistical Software, vol. 42 (6), pp. 1-29, 2012.

[Bouveyron et al., 2007] Bouveyron, C. Girard, S. and Schmid, C: High-Dimensional Data Clustering, Computational Statistics and Data Analysis, vol. 52 (1), pp. 502-519, 2007.

Examples

# Hepta
data("Hepta")
Data = Hepta$Data
#Non-default parameter model
#can be set to evaulate all possible models
V = HDDClustering(Data=Data,ClusterNo=7,model="ALL")
Cls = V$Cls

ClusterAccuracy(Hepta$Cls, Cls)

## Not run: 
library(HDclassif)
data(Crabs)
Data = Crabs[,-1]
V = HDDClustering(Data=Data,ClusterNo=4,com_dim=1)

## End(Not run)

FCPS documentation built on Oct. 3, 2026, 9:06 a.m.