Provides a flexible, simulation-based toolkit for exploring how
much data are needed to develop reliable prediction models. It works by
repeatedly generating data, fitting models, and evaluating performance
to show how sample size affects predictive accuracy, calibration, and
overfitting. The package supports continuous, binary, and time-to-event
outcomes and can be used with both regression-based modelling approaches
and machine-learning methods. It is designed to help researchers plan
studies, assess feasibility, and build more robust and generalisable
models. The methods are described in Olaniran et al. (2026)

pmsims is an R package for estimating how much data are needed to develop reliable and generalisable prediction models. It uses a simulation-based learning curve approach to quantify how model performance improves with increasing sample size, supporting principled study planning and feasibility assessment.
The package is fully model-agnostic: users can define how data are generated, how models are fitted, and how predictive performance is measured. Built-in workflows cover continuous, binary, and time-to-event outcomes, with a choice of regression-based models (linear, logistic, and Cox) and machine-learning models (regularised regression, random forest, and XGBoost).
Developed at King’s College London (Department of Biostatistics & Health Informatics) with input from researchers, clinicians, and patient partners. See the pmsims project site for further details.
Install the stable 1.0.0 release from GitHub:
# install.packages("remotes")
remotes::install_github("pmsims-package/pmsims", ref = "v1.0.0")
If you are interested in trying the development version, install from
the dev branch:
# install.packages("remotes")
remotes::install_github("pmsims-package/pmsims", ref = "dev")
The development version includes work in progress and may change before the next tagged release.
library(pmsims)
set.seed(123)
binary_example <- simulate_binary(
signal_parameters = 10,
noise_parameters = 10,
complexity = 2,
data_control = list(
nonlinear_strength = 0.4,
correlation = 0.2
),
outcome_prevalence = 0.20,
maximum_achievable_cstatistic = 0.75,
model = "glm",
metric = "calibration_slope",
target_performance = 0.90,
n_reps_total = 1000,
mean_or_assurance = "assurance"
)
binary_example
maximum_achievable_cstatistic and target_performance have different
roles:
maximum_achievable_cstatistic represents the best plausible
C-statistic with effectively unlimited data and calibrates the data
generator.target_performance is the minimum acceptable metric value used to
determine the required sample size.If you use pmsims, please cite the package and either or both
accompanying papers.
The validation paper:
The overview paper, currently a preprint:
Once the overview paper is published, that citation should be updated to the peer-reviewed version. In R, you can retrieve the package citation with:
citation("pmsims")
We welcome questions, suggestions, and collaboration enquiries.
This work is supported by the National Institute for Health and Care Research (NIHR) under the Research for Patient Benefit (RfPB) Programme (NIHR206858).
The views expressed are those of the authors and not necessarily those of the NIHR or the Department of Health and Social Care.