Create, Optimize, and Refine Data Nuggets

Creating, optimizing and refining data nuggets. Data nuggets reduce a large dataset into a small collection of nuggets of data, each containing a center (location), weight (importance), and scale (variability) parameter. Data nugget centers are selected based on a space-filling maximum-entropy scheme. Data nugget weights are created by counting the number observations closest to a given data nugget center. We then say the data nugget 'contains' these observations and the data nugget center is recalculated as the mean of these observations. Data nugget scales are created by calculating the trace of the covariance matrix of the observations contained within a data nugget divided by the dimension of the dataset. The optimal number of data nuggets is determined data-driven based on the relative second-order differences of propensity score indices. Data nuggets are refined by 'splitting' data nuggets which have high scales or elongated shapes (defined as the ratio of the two largest eigenvalues of the covariance matrix of the observations contained within the data nugget).


Reference manual

It appears you don't have a PDF plugin for this browser. You can click here to download the reference manual.

install.packages("datanugget")

1.5.0 by Rituparna Dey, a month ago


Browse source code at https://github.com/cran/datanugget


Authors: Rituparna Dey [aut, cre] , Yajie Duan [aut] , Traymon Beavers [aut] , Javier Cabrera [aut] , Ge Cheng [aut] , Kunting Qi [aut] , Mariusz Lubomirski [aut]


Documentation:   PDF Manual  


GPL-2 license


Depends on doSNOW, doParallel, foreach, parallel, Rfast, mgcv, ggplot2

Suggests testthat


Depended on by PPbigdata, WCluster.

Suggested by MosaiClusteR.


See at CRAN