Skip to main navigation Skip to search Skip to main content

Pretrained language models for selection of cohorts from electronic health records

Research output: Contribution to journalArticleAcademicpeer-review

Abstract

Clinical notes in Electronic Health Records (EHRs) can provide enormous amounts of data for research purposes. Extraction of data from clinical notes can be automated by Pretrained Language Models (PLMs). However, erroneous PLMs can extract incorrect data potentially impacting study results. We evaluated the impact of PLM-induced errors in cohort selection on subsequent clinical research with a focus on prognostic prediction model development. We used an EHR database of over 40,000 patients and deliberately decreased the performance of an PLM such that the model selected increasingly inaccurate cohorts of patients. Eligibility was defined by the presence of a target disease/procedure in the clinical notes. We used these inaccurate cohorts to develop prognostic prediction models and evaluated their discrimination and calibration performance. We found that PLMs with decreasing cohort selection performance (expressed by the F1-score), selected increasingly inaccurate cohorts. This resulted in aberrant regression coefficients in the prediction models developed on those cohorts. However, it did not affect the performance of the prediction models in terms of discrimination or calibration. While inaccurate cohort selections by erroneous PLMs affected the prediction model's coefficients, discrimination and calibration remained largely unaffected by PLM error. This is possibly due to the insensitivity of these measures to slight changes in the study cohort. Our results suggest that prognostic models may be robust to minor PLM-induced cohort selection errors. For clinical research focused on general prediction model performance rather than interpretation of individual coefficients, PLMs with moderate errors may still be applicable for selection of cohorts.

Original languageEnglish
Article number101803
JournalInformatics in Medicine Unlocked
Volume65
DOIs
Publication statusPublished - Sept 2026

Keywords

  • Electronic health records
  • Large Language Models
  • Natural Language Processing
  • Prediction modeling
  • Transformers

Fingerprint

Dive into the research topics of 'Pretrained language models for selection of cohorts from electronic health records'. Together they form a unique fingerprint.

Cite this