EVOKE: Elevating Chest X-ray Report Generation via Multi-View Contrastive Learning and Patient-Specific Knowledge

Qiguang Miao; Kang Liu; Zhuoqi Ma; Yunan Li; Xiaolu Kang; Ruixuan Liu; Tianyi Liu; Kun Xie; Zhicheng Jiao

EVOKE: Elevating Chest X-ray Report Generation via Multi-View Contrastive Learning and Patient-Specific Knowledge

Qiguang Miao, Kang Liu, Zhuoqi Ma, Yunan Li, Xiaolu Kang, Ruixuan Liu, Tianyi Liu, Kun Xie, Zhicheng Jiao

TL;DR

EVOKE tackles automated chest X-ray report generation by leveraging multi-view radiographs and patient INDICATION. It introduces a two-stage training framework: Stage 1 employs multi-view contrastive learning to align multiple views with report semantics, and Stage 2 performs knowledge-guided report generation using an INDICATION-aware transition bridge to handle missing indications. The approach achieves state-of-the-art results across MIMIC-CXR, MIMIC-ABN, Multi-view CXR, and Two-view CXR datasets, with significant improvements in RadGraph F1, BLEU, and CheXbert metrics, and demonstrates robustness through ablations and human evaluation. By providing curated Multi-view CXR and Two-view CXR datasets, EVOKE facilitates further research on multi-view radiology report generation and has potential to improve diagnostic efficiency and consistency in settings with diverse imaging views and available clinical indications.

Abstract

Radiology reports are crucial for planning treatment strategies and facilitating effective doctor-patient communication. However, the manual creation of these reports places a significant burden on radiologists. While automatic radiology report generation presents a promising solution, existing methods often rely on single-view radiographs, which constrain diagnostic accuracy. To address this challenge, we propose \textbf{EVOKE}, a novel chest X-ray report generation framework that incorporates multi-view contrastive learning and patient-specific knowledge. Specifically, we introduce a multi-view contrastive learning method that enhances visual representation by aligning multi-view radiographs with their corresponding report. After that, we present a knowledge-guided report generation module that integrates available patient-specific indications (e.g., symptom descriptions) to trigger the production of accurate and coherent radiology reports. To support research in multi-view report generation, we construct Multi-view CXR and Two-view CXR datasets using publicly available sources. Our proposed EVOKE surpasses recent state-of-the-art methods across multiple datasets, achieving a 2.9\% F\textsubscript{1} RadGraph improvement on MIMIC-CXR, a 7.3\% BLEU-1 improvement on MIMIC-ABN, a 3.1\% BLEU-4 improvement on Multi-view CXR, and an 8.2\% F\textsubscript{1,mic-14} CheXbert improvement on Two-view CXR.

EVOKE: Elevating Chest X-ray Report Generation via Multi-View Contrastive Learning and Patient-Specific Knowledge

TL;DR

Abstract

EVOKE: Elevating Chest X-ray Report Generation via Multi-View Contrastive Learning and Patient-Specific Knowledge

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (4)