cminkowski: Calculate the Minkowski distances for each pair of factors or...

View source: R/cmahalanobis.R

cminkowskiR Documentation

Calculate the Minkowski distances for each pair of factors or for the index.

Description

This function takes a dataframe and a variable or variables (two or more) in input, and returns a matrix or matrices (two or more) with the Minkowski distances about the factors inside them. You can also select "index" to calculate the Minkowski distances between each row.

Usage

cminkowski(
  dataset,
  formula,
  p = 3,
  plot = TRUE,
  min_group_size = 3,
  method = "average",
  max_index_sample = NULL,
  automatic_encoding = FALSE,
  na_removal = FALSE,
  grouping_stat = "mean"
)

Arguments

dataset

A dataframe.

formula

The index of the dataframe, otherwise a variable or variables (two or more) with factors which you want to calculate the Minkowski distances matrix or matrices (two or more).

p

Order of the Minkowski distance.

plot

Logical, if TRUE, dendrograms with various agglomeration metrics for factors (two or more) are displayed. With "index" in formula, a dendrogram considering the observation is displayed.

min_group_size

Minimum group size to maintain. The default value is 3, therefore groups, inside variables, with less than 3 observations will be discarded. For "index", this value is always 1.

method

The agglomeration method for calculating dissimilarities between observations. Available methods are "ward.D", "ward.D2", "single", "complete", "average", "mcquitty", "median" or "centroid".

max_index_sample

A number of random samples from a dataset for which you want to calculate the dissimilarity for index mode, useful with very large dataset.

automatic_encoding

Logical, if TRUE, names inside factor variables will be transformed in numbers with ordinal order (1,2,....).

na_removal

Logical, if TRUE, missing value removal on rows is performed.

grouping_stat

When a factor variable is specified, calculate the specified grouping statistic for each factor. Available methods are: mean (arithmetic mean), median and SDS (standard deviations).

Value

According to the option chosen in formula, with "index" the Minkowski distances matrix will be printed; instead, by specifying variables, the Minkowski distances matrix or matrices (two or more) between each pair of groups and, optionally, the plot or plots (two or more) will be printed.

Note

When p < 1, the Minkowski distance is a "dissimilarity" measure. When p >= 1, the triangle inequality property is satisfied and we say "Minkowski distance". If "index" is selected with variables, only distances between rows are calculated. Therefore, this snippet: "cminkowski(mtcars, ~am + carb + index)" will print distances only considering "index". Optionally, rows with NA values are omitted.

Examples


# Example with the CO2 dataset

table(CO2$Plant)

cminkowski(CO2, ~Plant, p = 4, 
              plot = TRUE, 
              grouping_stat = 'mean', 
              na_removal = TRUE, 
              automatic_encoding = TRUE)

# Example with the airquality dataset

summary(airquality)

cminkowski(airquality, ~index, p = 3,
   plot = TRUE, 
   na_removal = TRUE, 
   max_index_sample = 40)


cmahalanobis documentation built on Aug. 31, 2026, 5:07 p.m.