Large Language Models are Powerful Electronic Health Record Encoders

Stefan Hegselmann; Georg von Arnim; Tillmann Rheude; Noel Kronenberg; David Sontag; Gerhard Hindricks; Roland Eils; Benjamin Wild

Large Language Models are Powerful Electronic Health Record Encoders

Stefan Hegselmann, Georg von Arnim, Tillmann Rheude, Noel Kronenberg, David Sontag, Gerhard Hindricks, Roland Eils, Benjamin Wild

TL;DR

This work investigates repurposing general-purpose Large Language Models (LLMs) to encode Electronic Health Records (EHRs) by serializing structured records into text, enabling high-dimensional embeddings that drive clinical predictions without private, institution-specific training data. The approach is evaluated on the EHRSHOT benchmark and an external UK Biobank cohort, showing that LLM embeddings can match or surpass a domain-specific EHR foundation model (CLMBR-T-Base) across multiple tasks, especially under domain shifts and low-data regimes. Key findings include the robustness of Markdown-style EHR serialization across formats, the benefit of focusing on recent history with extended context windows, and the complementary potential of combining LLM embeddings with domain-specific representations. The results argue for scalable, interoperable EHR encoders that leverage broad text pretraining, enabling cross-institution applicability and flexible integration with standard discriminative heads, while highlighting tradeoffs in computation, calibration, and the need for broader external validation.

Abstract

Electronic Health Records (EHRs) offer considerable potential for clinical prediction, but their complexity and heterogeneity present significant challenges for traditional machine learning methods. Recently, domain-specific EHR foundation models trained on large volumes of unlabeled EHR data have shown improved predictive accuracy and generalization. However, their development is constrained by limited access to diverse, high-quality datasets, and inconsistencies in coding standards and clinical practices. In this study, we explore the use of general-purpose Large Language Models (LLMs) to encode EHR into high-dimensional representations for downstream clinical prediction tasks. We convert structured EHR data into Markdown-formatted plain-text documents by replacing medical codes with natural language descriptions. This enables the use of LLMs and their extensive semantic understanding and generalization capabilities as effective encoders of EHRs without requiring access to private medical training data. We show that LLM-based embeddings can often match or even surpass the performance of a specialized EHR foundation model, CLMBR-T-Base, across 15 diverse clinical tasks from the EHRSHOT benchmark. Critically, our approach requires no institution-specific training and can incorporate any medical code with a text description, whereas existing EHR foundation models operate on fixed vocabularies and can only process codes seen during pretraining. To demonstrate generalizability, we further evaluate the approach on the UK Biobank (UKB) cohort, out-of-domain for CLMBR-T-Base, whose fixed vocabulary covers only 16% of UKB codes. Notably, an LLM-based model achieves superior performance for prediction of disease onset, hospitalization, and mortality, indicating robustness to population and coding shifts.

Large Language Models are Powerful Electronic Health Record Encoders

TL;DR

Abstract

Large Language Models are Powerful Electronic Health Record Encoders

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (17)