| cjaccard | R Documentation |
This function takes a dataframe and a variable or variables (two or more) in input, and returns a matrix or matrices (two or more) with the Jaccard distance about the factors inside them. You can also select "index" to calculate the Jaccard distances between each row.
cjaccard(
dataset,
formula,
plot = TRUE,
min_group_size = 3,
method = "average",
max_index_sample = NULL,
automatic_encoding = FALSE,
na_removal = FALSE,
grouping_stat = "mean"
)
dataset |
A dataframe. |
formula |
The index of the dataframe, otherwise a variable or variables (two or more) with factors which you want to calculate the Jaccard distances matrix or matrices (two or more). |
plot |
Logical, if TRUE, dendrograms with various agglomeration metrics for factors (two or more) are displayed. With "index" in formula, a dendrogram considering the observation is displayed. |
min_group_size |
Minimum group size to maintain. The default value is 3, therefore groups, inside variables, with less than 3 observations will be discarded. For "index", this value is always 1. |
method |
The agglomeration method for calculating distances between observations. Available methods are "median" or "centroid". |
max_index_sample |
A number of random samples from a dataset for which you want to calculate the distance for index mode, useful with very large dataset. |
automatic_encoding |
Logical, if TRUE, names inside factor variables will be transformed in numbers with ordinal order (1,2,....). |
na_removal |
Logical, if TRUE, missing value removal on rows is performed. |
grouping_stat |
When a factor variable is specified, calculate the specified grouping statistic for each factor. Available methods are: mean (arithmetic mean), median and SDS (standard deviations). |
According to the option chosen in formula, with "index" the Jaccard distances matrix will be printed; instead, by specifying variables, the Jaccard distances matrix or matrices (two or more) between each pair of groups and, optionally, the plot or plots (two or more) will be printed.
If "index" is selected with variables, only dissimilarities between rows are calculated. Therefore, this snippet: "cjaccard(mtcars, ~am + carb + index)" will print dissimilarities only considering "index". Optionally, rows with NA values are omitted.
# Example with the CO2 dataset
table(CO2$Plant)
cjaccard(CO2, ~Plant,
plot = TRUE,
grouping_stat = 'mean',
na_removal = TRUE,
automatic_encoding = TRUE)
# Example with the airquality dataset
summary(airquality)
cjaccard(airquality, ~index,
plot = TRUE,
na_removal = TRUE,
max_index_sample = 40)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.