TY - JOUR
T1 - Internal-external cross-validation helped to evaluate the generalizability of prediction models in large clustered datasets
AU - Takada, Toshihiko
AU - Nijman, Steven
AU - Denaxas, Spiros
AU - Snell, Kym I E
AU - Uijl, Alicia
AU - Nguyen, Tri-Long
AU - Asselbergs, Folkert W
AU - Debray, Thomas P A
N1 - Funding Information:
This work has received support from the EU/EFPIA Innovative Medicines Initiative 2 Joint Undertaking BigData@Heart grant no. 116074 . TT was supported by the Uehara Memorial Foundation between August 2018 and July 2019. SN is supported by a Public-Private Study grant of the Dutch Heart foundation ( 2018B006 ). KIES is supported by a National Institute for Health Research School for Primary Care Research (NIHR SPCR) Launching Fellowship. FWA is supported by UCL Hospitals NIHR Biomedical Research Centre . TPAD is supported by the Netherlands Organisation for Health Research and Development (91617050) , and by the European Union's Horizon 2020 research and innovation programme under ReCoDID grant agreement no. 825746.
Funding Information:
Case study 1: Due to privacy laws and the data user agreement between the University College London and Clinical Practice Research Datalink, authors are not authorised to share individual patient data from these electronic health records. Requests to access data provided by Clinical Practice Research Datalink (CPRD) should be sent to the Independent Scientific Advisory Committee (ISAC) (https://www.cprd. com/ISAC/). The CALIBER portal (https://www.caliberresearch.org/portal/) does offer open sharing of phenotypic and analytic algorithms for use by other researchers. Aggregate data (e.g. model formulas, performance estimates) are available on request. Case study 2: The International Stroke Trial database is available from https://datashare.ed.ac.uk/handle/10283/124. This work has received support from the EU/EFPIA Innovative Medicines Initiative 2 Joint Undertaking BigData@Heart grant no. 116074. TT was supported by the Uehara Memorial Foundation between August 2018 and July 2019. SN is supported by a Public-Private Study grant of the Dutch Heart foundation (2018B006). KIES is supported by a National Institute for Health Research School for Primary Care Research (NIHR SPCR) Launching Fellowship. FWA is supported by UCL Hospitals NIHR Biomedical Research Centre. TPAD is supported by the Netherlands Organisation for Health Research and Development (91617050), and by the European Union's Horizon 2020 research and innovation programme under ReCoDID grant agreement no. 825746.[Formula presented], The views expressed are those of the author(s) and not necessarily those of the NIHR or the Department of Health and Social Care. This study is based in part on data from the Clinical Practice Research Datalink obtained under licence from the UK Medicines and Healthcare products Regulatory Agency. The data is provided by patients and collected by the NHS as part of their care and support. The interpretation and conclusions contained in this study are those of the author/s alone. Copyright ? (2020), re-used with the permission of The Health & Social Care Information Centre. All rights reserved.
Publisher Copyright:
© 2021 The Authors
PY - 2021/9
Y1 - 2021/9
N2 - OBJECTIVE: To illustrate how to evaluate the need of complex strategies for developing generalizable prediction models in large clustered datasets.STUDY DESIGN AND SETTING: We developed eight Cox regression models to estimate the risk of heart failure using a large population-level dataset. These models differed in the number of predictors, the functional form of the predictor effects (non-linear effects and interaction) and the estimation method (maximum likelihood and penalization). Internal-external cross-validation was used to evaluate the models' generalizability across the included general practices.RESULTS: Among 871,687 individuals from 225 general practices, 43,987 (5.5%) developed heart failure during a median follow-up time of 5.8 years. For discrimination, the simplest prediction model yielded a good concordance statistic, which was not much improved by adopting complex strategies. Between-practice heterogeneity in discrimination was similar in all models. For calibration, the simplest model performed satisfactorily. Although accounting for non-linear effects and interaction slightly improved the calibration slope, it also led to more heterogeneity in the observed/expected ratio. Similar results were found in a second case study involving patients with stroke.CONCLUSION: In large clustered datasets, prediction model studies may adopt internal-external cross-validation to evaluate the generalizability of competing models, and to identify promising modelling strategies.
AB - OBJECTIVE: To illustrate how to evaluate the need of complex strategies for developing generalizable prediction models in large clustered datasets.STUDY DESIGN AND SETTING: We developed eight Cox regression models to estimate the risk of heart failure using a large population-level dataset. These models differed in the number of predictors, the functional form of the predictor effects (non-linear effects and interaction) and the estimation method (maximum likelihood and penalization). Internal-external cross-validation was used to evaluate the models' generalizability across the included general practices.RESULTS: Among 871,687 individuals from 225 general practices, 43,987 (5.5%) developed heart failure during a median follow-up time of 5.8 years. For discrimination, the simplest prediction model yielded a good concordance statistic, which was not much improved by adopting complex strategies. Between-practice heterogeneity in discrimination was similar in all models. For calibration, the simplest model performed satisfactorily. Although accounting for non-linear effects and interaction slightly improved the calibration slope, it also led to more heterogeneity in the observed/expected ratio. Similar results were found in a second case study involving patients with stroke.CONCLUSION: In large clustered datasets, prediction model studies may adopt internal-external cross-validation to evaluate the generalizability of competing models, and to identify promising modelling strategies.
KW - Calibration
KW - Discrimination
KW - Heterogeneity
KW - Model comparison
KW - Prediction model
KW - Validation
UR - https://www.scopus.com/pages/publications/85105014589
U2 - 10.1016/j.jclinepi.2021.03.025
DO - 10.1016/j.jclinepi.2021.03.025
M3 - Article
C2 - 33836256
SN - 0895-4356
VL - 137
SP - 83
EP - 91
JO - Journal of Clinical Epidemiology
JF - Journal of Clinical Epidemiology
ER -