Examples: visualization, C++, networks, data cleaning, html widgets, ropensci.

Found 2116 packages in 0.03 seconds

rfPermute — by Eric Archer, a year ago

Estimate Permutation p-Values for Random Forest Importance Metrics

Estimate significance of importance metrics for a Random Forest model by permuting the response variable. Produces null distribution of importance metrics for each predictor variable and p-value of observed. Provides summary and visualization functions for 'randomForest' results.

randomForestVIP — by Kelvyn Bladen, 3 years ago

Tune Random Forests Based on Variable Importance & Plot Results

Functions for assessing variable relations and associations prior to modeling with a Random Forest algorithm (although these are relevant for any predictive model). Metrics such as partial correlations and variance inflation factors are tabulated as well as plotted for the user. A function is available for tuning the main Random Forest hyper-parameter based on model performance and variable importance metrics. This grid-search technique provides tables and plots showing the effect of the main hyper-parameter on each of the assessment metrics. It also returns each of the evaluated models to the user. The package also provides superior variable importance plots for individual models. All of the plots are developed so that the user has the ability to edit and improve further upon the plots. Derivations and methodology are described in Bladen (2022) < https://digitalcommons.usu.edu/etd/8587/>.

rQSAR — by Oche Ambrose George, 2 years ago

QSAR Modeling with Multiple Algorithms: MLR, PLS, and Random Forest

Quantitative Structure-Activity Relationship (QSAR) modeling is a valuable tool in computational chemistry and drug design, where it aims to predict the activity or property of chemical compounds based on their molecular structure. In this vignette, we present the 'rQSAR' package, which provides functions for variable selection and QSAR modeling using Multiple Linear Regression (MLR), Partial Least Squares (PLS), and Random Forest algorithms.

PowerXgammaRF — by Shikhar Tyagi, a month ago

Random Forest Regression with Power Xgamma Distribution Error Model

Implements Random Forest regression under the Power Xgamma distribution error model. Provides core distribution functions (density, cumulative distribution, quantile, random generation, hazard, survival), parameter estimation via Expectation-Maximization (EM) and Markov Chain Monte Carlo (MCMC), non-parametric bootstrap confidence intervals (at 90%, 95%, and 99% levels), Highest Posterior Density (HPD) intervals, Heidelberger and Welch's MCMC convergence diagnostic, convergence probability, model evaluation metrics (estimated values, bias, mean squared error, risk value), homoscedastic prediction intervals, and goodness-of-fit diagnostic tests (Kolmogorov-Smirnov and Anderson-Darling tests, Akaike Information Criterion, and Bayesian Information Criterion). References: Tyagi et al. (2022, Int. J. Stat. Reliab. Eng., 9(1), 51-60); Breiman (2001) ; Wright and Ziegler (2017) ; Heidelberger and Welch (1983) ; Sen et al. (2016) .

RFmerge — by Mauricio Zambrano-Bigiarini, 5 months ago

Merging of Satellite Datasets with Ground Observations using Random Forests

S3 implementation of the Random Forest MErging Procedure (RF-MEP), which combines two or more satellite-based datasets (e.g., precipitation products, topography) with ground observations to produce a new dataset with improved spatio-temporal distribution of the target field. In particular, this package was developed to merge different Satellite-based Rainfall Estimates (SREs) with measurements from rain gauges, in order to obtain a new precipitation dataset where the time series in the rain gauges are used to correct different types of errors present in the SREs. However, this package might be used to merge other hydrological/environmental gridded datasets with point observations. For details, see Baez-Villanueva et al. (2020) . Bugs / comments / questions / collaboration of any kind are very welcomed.

MTLRF — by Shikhar Tyagi, a month ago

Random Forest Regression with Modified Topp-Leone Error Model

Implements Random Forest regression under the Modified Topp-Leone (MTL) distribution error model. Provides core distribution functions (density, cumulative distribution, exact closed-form quantile, random generation, hazard, and survival), parameter estimation via closed-form Expectation-Maximization/Maximum Likelihood (EM/MLE) and Bayesian Markov Chain Monte Carlo (MCMC), non-parametric bootstrap confidence intervals (at 90%, 95%, and 99% levels), Highest Posterior Density (HPD) intervals, Heidelberger and Welch MCMC convergence diagnostics, model evaluation metrics (estimated values, bias, mean squared error, risk value), homoscedastic prediction intervals, and goodness-of-fit diagnostic tests (Kolmogorov-Smirnov and Anderson-Darling tests, Akaike Information Criterion, and Bayesian Information Criterion). References: Breiman (2001) ; Singh, Tyagi, Singh, and Tyagi (2025) < https://statassoc.or.th>; Topp and Leone (1955) ; Wright and Ziegler (2017) ; Plummer, Best, Cowles, and Vines (2006) < https://CRAN.R-project.org/package=coda>; Heidelberger and Welch (1983) .

missForestPredict — by Elena Albu, a year ago

Missing Value Imputation using Random Forest for Prediction Settings

Missing data imputation based on the 'missForest' algorithm (Stekhoven, Daniel J (2012) ) with adaptations for prediction settings. The function missForest() is used to impute a (training) dataset with missing values and to learn imputation models that can be later used for imputing new observations. The function missForestPredict() is used to impute one or multiple new observations (test set) using the models learned on the training data. For more details see Albu, E., Gao, S., Wynants, L., & Van Calster, B. (2024). missForestPredict--Missing data imputation for prediction settings .

forestControl — by Tom Wilson, 5 years ago

Approximate False Positive Rate Control in Selection Frequency for Random Forest

Approximate false positive rate control in selection frequency for random forest using the methods described by Ender Konukoglu and Melanie Ganz (2014) . Methods for calculating the selection frequency threshold at false positive rates and selection frequency false positive rate feature selection.

ModelMap — by Elizabeth Freeman, a year ago

Modeling and Map Production using Random Forest and Related Stochastic Models

Creates sophisticated models of training data and validates the models with an independent test set, cross validation, or Out Of Bag (OOB) predictions on the training data. Create graphs and tables of the model validation results. Applies these models to GIS .img files of predictors to create detailed prediction surfaces. Handles large predictor files for map making, by reading in the .img files in chunks, and output to the .txt file the prediction for each data chunk, before reading the next chunk of data.

JRF — by Francesca Petralia Developer, 10 years ago

Joint Random Forest (JRF) for the Simultaneous Estimation of Multiple Related Networks

Simultaneous estimation of multiple related networks.