optimal.DN: Optimal number of Data Nuggets

View source: R/optimal.DN.R

optimal.DNR Documentation

Optimal number of Data Nuggets

Description

This function finds the optimal number data nuggets, and creates the nuggets using the create.DN function.

Usage

optimal.DN(x,
           center.method = "mean", 
           dn.nos, 
           R = 5000,
           delete.percent = .1,
           DN.num1 = 10^4,
           eps = 5e-3,
           dist.metric = "euclidean", 
           seed = 291102,
           no.cores = (parallel::detectCores() - 1),
           make.pbs = FALSE)

Arguments

x

A data matrix (of class matrix, data.frame, or data.table) containing only entries of class numeric.

center.method

The method used for creating data nugget centers. Must be 'mean' or 'random' or 'original'. 'mean' chooses the data nugget center to be the mean of all observations within that data nugget, 'random' chooses the data nugget center to be some random observation within that data nugget, and 'original' chooses the original data nugget centers generated by the final run of datanugget creation using create.DNcenters function. Default is 'mean'.

dn.nos

The vector of candidate datanugget numbers. Must be a vector of length >= 3 and have numeric or integer entries.

R

The number of observations to sample from the data matrix when creating the initial data nugget centers. Must be of class numeric within [100,10000]. Default is 5000.

delete.percent

The proportion of observations to remove from the data matrix at each iteration when finding data nugget centers. Must be of class numeric and within (0,1). Default is 0.1.

DN.num1

The number of initial data nugget centers to create. Must be of class numeric. Default is 10^4.

eps

Stoppage tolerance on hitting the elbow. Default is 5e-3.

dist.metric

The distance metric used to create the initial centers of data nuggets. Must be 'euclidean' or 'manhattan'. Default is 'euclidean'.

seed

Random seed for replication. Must be of class numeric. Default is 291102.

no.cores

Number of cores used for parallel processing. If '0' then parallel processing is not used. Must be of class numeric.

make.pbs

Logical; whether to show a progress bar while the function runs. Default is FALSE.

Details

The optimal number data nuggets is data-driven based on the relative second-order differences of propensity score indices. That optimal number of data nuggets are created using the create.DN function.

Value

A list of 4 items:

opt.dn.no

Optimal data nugget number.

opt.dn

Final datanugget object based on optimal data nugget number.

elbow.plot

Elbow plot of Propensity Score Index vs Data Nugget number.

diff.plot

Relative second order differences vs Data Nugget number plot.

Author(s)

Rituparna Dey, Javier Cabrera

References

Beavers, T. E., Cheng, G., Duan, Y., Cabrera, J., Lubomirski, M., Amaratunga, D., & Teigler, J. E. (2024). Data Nuggets: A Method for Reducing Big Data While Preserving Data Structure. Journal of Computational and Graphical Statistics, 1-21.

Cherasia, K. E., Cabrera, J., Fernholz, L. T., & Fernholz, R. (2022). Data Nuggets in Supervised Learning. In Robust and Multivariate Statistical Methods: Festschrift in Honor of David E. Tyler (pp. 429-449). Cham: Springer International Publishing.

Examples


      ## small example
      X = cbind.data.frame(rnorm(10^3),
                           rnorm(10^3),
                           rnorm(10^3))

      suppressMessages({

        my.DN = optimal.DN(x = X,
                           dn.nos = seq(50, 500, by = 25),
                           R = 500,
                           delete.percent = .1,
                           DN.num1 = 500,
                           eps = 5e-5,
                           no.cores = 0,
                           make.pbs = FALSE)

      })

      my.DN$opt.dn.no
      my.DN$opt.dn
      my.DN$elbow.plot
      my.DN$diff.plot

    ## Not run: 

      ## large example
      X = cbind.data.frame(rnorm(5*10^4),
                           rnorm(5*10^4),
                           rnorm(5*10^4),
                           rnorm(5*10^4),
                           rnorm(5*10^4))
      t1 <- Sys.time()
      my.DN = optimal.DN(x = X,
                         dn.nos = seq(1000, 10000, by = 1000),
                         R = 5000,
                         delete.percent = .1,
                         DN.num1 = 10^4,
                         eps = 5e-5,
                         no.cores = 2)
      t2 <- Sys.time()
      
      my.DN$opt.dn.no
      my.DN$opt.dn
      my.DN$elbow.plot
      my.DN$diff.plot

    
## End(Not run)


datanugget documentation built on Aug. 21, 2026, 9:10 a.m.