TY - GEN
T1 - Estimating the Sample Size Required to Minimize Model Uncertainty Across Different Risk Profiles
AU - Vereijken, Fleur
AU - Reps, Jenna M.
AU - Fridgeirsson, Egill A.
AU - Hackmann, Toby
AU - Steyerberg, Ewout W.
AU - Rijnbeek, Peter
AU - Williams, Ross D.
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/6/25
Y1 - 2026/6/25
N2 - Clinical prediction models are increasingly used in healthcare to estimate an individual's risk of a future outcome, such as 10-year cardiovascular risk in primary care populations. It is often assumed that a well-defined prediction model provides a single, adequate estimate for a given prediction task and dataset. This, however, overlooks that equally valid models can be derived from the same underlying data which assign quite different predictions to the same individual. This model uncertainty makes model evaluation and interpretation challenging. In this paper we studied stability of model performance and individual-level predictions as a function of increasing model development sample size. We hypothesize that model performance and individual predictions will stabilize with more development data, but individual predictions will require more data to stabilize. We varied the training sample size by increasing the number of outcome events included in the training data, ranging from 350 to 5,000 events (if available), while preserving the original event-to-non-event ratio. Models were trained using age + sex only as predictors in comparison to using predictor sets without restrictions, with several hundred predictors. With increasing sample size, discrimination increased, while variability in estimated probabilities within individuals decreased. Instability was greatest among higher-risk individuals on the probability scale, while on the logit scale, slightly greater variability was observed among lowest-risk individuals. Extreme profiles thus have relatively large uncertainty. The severity of class imbalance impacted the severity of individual instability on the probability scale, but not on the logit scale. Not only predicted probabilities for individuals were unstable, but the relative ranking of individual patients also differed between models. Increasing the number of predictors often improved discriminative performance but was associated with greater instability in individual predictions, whereas simpler models with fewer predictors showed more stable predictions. Although increasing sample size reduced variability in model predictions, instability did not disappear entirely. In conclusion, model uncertainty depends on sample size and flexibility of prediction models. It highlights a trade-off between predictive performance and stability that should be considered when developing and evaluating prediction models.
AB - Clinical prediction models are increasingly used in healthcare to estimate an individual's risk of a future outcome, such as 10-year cardiovascular risk in primary care populations. It is often assumed that a well-defined prediction model provides a single, adequate estimate for a given prediction task and dataset. This, however, overlooks that equally valid models can be derived from the same underlying data which assign quite different predictions to the same individual. This model uncertainty makes model evaluation and interpretation challenging. In this paper we studied stability of model performance and individual-level predictions as a function of increasing model development sample size. We hypothesize that model performance and individual predictions will stabilize with more development data, but individual predictions will require more data to stabilize. We varied the training sample size by increasing the number of outcome events included in the training data, ranging from 350 to 5,000 events (if available), while preserving the original event-to-non-event ratio. Models were trained using age + sex only as predictors in comparison to using predictor sets without restrictions, with several hundred predictors. With increasing sample size, discrimination increased, while variability in estimated probabilities within individuals decreased. Instability was greatest among higher-risk individuals on the probability scale, while on the logit scale, slightly greater variability was observed among lowest-risk individuals. Extreme profiles thus have relatively large uncertainty. The severity of class imbalance impacted the severity of individual instability on the probability scale, but not on the logit scale. Not only predicted probabilities for individuals were unstable, but the relative ranking of individual patients also differed between models. Increasing the number of predictors often improved discriminative performance but was associated with greater instability in individual predictions, whereas simpler models with fewer predictors showed more stable predictions. Although increasing sample size reduced variability in model predictions, instability did not disappear entirely. In conclusion, model uncertainty depends on sample size and flexibility of prediction models. It highlights a trade-off between predictive performance and stability that should be considered when developing and evaluating prediction models.
KW - learning curves
KW - Machine learning
KW - Model uncertainty
UR - https://www.scopus.com/pages/publications/105044415869
U2 - 10.1145/3805689.3812316
DO - 10.1145/3805689.3812316
M3 - Conference contribution
AN - SCOPUS:105044415869
T3 - ACM FAccT 2026 - Proceedings of the 9th annual ACM Conference on Fairness, Accountability, and Transparency
SP - 3302
EP - 3321
BT - ACM FAccT 2026 - Proceedings of the 9th annual ACM Conference on Fairness, Accountability, and Transparency
PB - Association for Computing Machinery
T2 - 9th Annual ACM Conference on Fairness, Accountability, and Transparency, ACM FAccT 2026
Y2 - 25 June 2026 through 28 June 2026
ER -