| refine.DN | R Documentation |
This function refines the data nuggets found in an object of class datanugget created using the create.DN function or optimal.DN function.
refine.DN(x,
DN,
EV.tol = 0.9,
max.splits = 5,
min.nugget.size = 2,
shape.split = FALSE,
min.shape.size = 10,
delta = 2,
nstart = 25,
seed = 291102,
no.cores = (parallel::detectCores() - 1),
make.pbs = FALSE)
x |
A data matrix (data frame, data table, matrix, etc.) containing only entries of class numeric. |
DN |
An object of class data nugget created using the create.DN or optimal.DN function. |
EV.tol |
A value designating the percentile for finding the corresponding quantile that will designate how large the largest eigenvalue of the covariance matrix of a data nugget can be before it must be split. Must be of class numeric and within (0,1). Default is 0.9. |
max.splits |
A value designating the maximum amount of attempts that will be made to split data nuggets according to their largest eigenvalue before the algorithm breaks. Must be of class numeric and non-negative. Default is 5. |
min.nugget.size |
A value designating the minimum amount of observations a data nugget created from a split must contain. Must be of class numeric and with be greater than 1. Default is 2. |
shape.split |
Logical; whether to refine the nuggets of elongated shape i.e., perform shape splitting. Default is FALSE. |
min.shape.size |
A value designating the minimum weight for a data nugget to be shape split. Must be at least 10. Default is 10. |
delta |
A value defining a data nugget to be elongated when its first two largest eigenvalues, EV1 and EV2 satisfy EV1/EV2 > delta. Must be gtreater than 1. Default is 2. |
nstart |
The number of random starts used by the kmeans algorithm when splitting. Default is 25. |
seed |
Random seed for replication. Must be of class numeric. Default is 291102. |
no.cores |
Number of cores used for parallel processing. If '0' then parallel processing is not used. Must be of class numeric. |
make.pbs |
Print progress bars? Must be TRUE or FALSE. |
Data nuggets can be refined by splitting the nuggets with unusually large within-data-nugget variability. For each nugget, the largest eigenvalue of its covariance matrix, EV.1 serves as the variability measure. EV.tol defines the percentile threshold and if a nugget's EV.1 exceeds that threshold, it is split into two by K-Means algorithm. However, if either of the two data nuggets created by this split have less than the designated minimum data nugget size (min.nugget.size), then the split is cancelled and the data nugget remains as is. This function refines data nuggets using Algorithm 2 provided in the reference.
Another version of refinement is shape based splitting. A data nugget is considered to be very elongated in shape, if the ratio of the first to the second largest eigenvalues of its covariance matrix, EV.1 / EV.2 is greater than a specified threshold, delta (default value is 2). Those data nuggets are split using the same K-Means clustering.
An object of class datanugget:
Data Nuggets |
|
Data Nugget Assignments |
Vector of length |
Rituparna Dey, Traymon Beavers, Javier Cabrera, Mariusz Lubomirski
Beavers, T. E., Cheng, G., Duan, Y., Cabrera, J., Lubomirski, M., Amaratunga, D., & Teigler, J. E. (2024). Data Nuggets: A Method for Reducing Big Data While Preserving Data Structure. Journal of Computational and Graphical Statistics, 1-21.
Cherasia, K. E., Cabrera, J., Fernholz, L. T., & Fernholz, R. (2022). Data Nuggets in Supervised Learning. In Robust and Multivariate Statistical Methods: Festschrift in Honor of David E. Tyler (pp. 429-449). Cham: Springer International Publishing.
## small example
X = cbind.data.frame(rnorm(10^3),
rnorm(10^3),
rnorm(10^3))
suppressMessages({
my.DN = create.DN(x = X,
R = 500,
delete.percent = .1,
DN.num1 = 500,
DN.num2 = 250,
no.cores = 0,
make.pbs = FALSE)
my.DN2 = refine.DN(x = X,
DN = my.DN,
no.cores = 0,
make.pbs = FALSE)
})
my.DN2$`Data Nuggets`
my.DN2$`Data Nugget Assignments`
## Not run:
## large example
X = cbind.data.frame(rnorm(5*10^4),
rnorm(5*10^4),
rnorm(5*10^4),
rnorm(5*10^4),
rnorm(5*10^4))
my.DN = create.DN(x = X,
R = 5000,
delete.percent = .1,
DN.num1 = 10^4,
DN.num2 = 2000,
no.cores = 2)
my.DN2 = refine.DN(x = X,
DN = my.DN,
no.cores = 2)
my.DN2$`Data Nuggets`
my.DN2$`Data Nugget Assignments`
## End(Not run)
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.