Skip to main navigation Skip to search Skip to main content

Estimating the Sample Size Required to Minimize Model Uncertainty Across Different Risk Profiles

  • Fleur Vereijken*
  • , Jenna M. Reps
  • , Egill A. Fridgeirsson
  • , Toby Hackmann
  • , Ewout W. Steyerberg
  • , Peter Rijnbeek
  • , Ross D. Williams
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contributionAcademicpeer-review

Abstract

Clinical prediction models are increasingly used in healthcare to estimate an individual's risk of a future outcome, such as 10-year cardiovascular risk in primary care populations. It is often assumed that a well-defined prediction model provides a single, adequate estimate for a given prediction task and dataset. This, however, overlooks that equally valid models can be derived from the same underlying data which assign quite different predictions to the same individual. This model uncertainty makes model evaluation and interpretation challenging. In this paper we studied stability of model performance and individual-level predictions as a function of increasing model development sample size. We hypothesize that model performance and individual predictions will stabilize with more development data, but individual predictions will require more data to stabilize. We varied the training sample size by increasing the number of outcome events included in the training data, ranging from 350 to 5,000 events (if available), while preserving the original event-to-non-event ratio. Models were trained using age + sex only as predictors in comparison to using predictor sets without restrictions, with several hundred predictors. With increasing sample size, discrimination increased, while variability in estimated probabilities within individuals decreased. Instability was greatest among higher-risk individuals on the probability scale, while on the logit scale, slightly greater variability was observed among lowest-risk individuals. Extreme profiles thus have relatively large uncertainty. The severity of class imbalance impacted the severity of individual instability on the probability scale, but not on the logit scale. Not only predicted probabilities for individuals were unstable, but the relative ranking of individual patients also differed between models. Increasing the number of predictors often improved discriminative performance but was associated with greater instability in individual predictions, whereas simpler models with fewer predictors showed more stable predictions. Although increasing sample size reduced variability in model predictions, instability did not disappear entirely. In conclusion, model uncertainty depends on sample size and flexibility of prediction models. It highlights a trade-off between predictive performance and stability that should be considered when developing and evaluating prediction models.

Original languageEnglish
Title of host publicationACM FAccT 2026 - Proceedings of the 9th annual ACM Conference on Fairness, Accountability, and Transparency
PublisherAssociation for Computing Machinery
Pages3302-3321
Number of pages20
ISBN (Electronic)979-8-4007-2596-8
DOIs
Publication statusPublished - 25 Jun 2026
Event9th Annual ACM Conference on Fairness, Accountability, and Transparency, ACM FAccT 2026 - Montreal, Canada
Duration: 25 Jun 202628 Jun 2026

Publication series

NameACM FAccT 2026 - Proceedings of the 9th annual ACM Conference on Fairness, Accountability, and Transparency

Conference

Conference9th Annual ACM Conference on Fairness, Accountability, and Transparency, ACM FAccT 2026
Country/TerritoryCanada
CityMontreal
Period25/06/2628/06/26

Keywords

  • learning curves
  • Machine learning
  • Model uncertainty

Fingerprint

Dive into the research topics of 'Estimating the Sample Size Required to Minimize Model Uncertainty Across Different Risk Profiles'. Together they form a unique fingerprint.

Cite this