On the Use of Large Language Models for Qualitative Synthesis
Sebastián Pizard, Ramiro Moreira, Federico Galiano, Ignacio Sastre, Lorena Etcheverry
TL;DR
This paper investigates the challenges of using large language models to support qualitative synthesis within systematic reviews. It employs a collaborative autoethnographic approach across two trials to characterize how QS can be supported or hindered by LLMs, identifying 13 key challenges related to nuance, evaluation, bias, transparency, and reproducibility. The authors propose a set of QS-specific characteristics and evaluative criteria to judge soundness, and advocate for risk management and human-in-the-loop oversight, guided by the RAISE framework. The work highlights that while LLMs can aid narrow tasks such as descriptive categorization, achieving a rigorous and trustworthy QS often demands substantial effort comparable to manual methods, underscoring the need for cautious deployment and further research, especially with open-source or task-specific tools.
Abstract
Large language models (LLMs) show promise for supporting systematic reviews (SR), even complex tasks such as qualitative synthesis (QS). However, applying them to a stage that is unevenly reported and variably conducted carries important risks: misuse can amplify existing weaknesses and erode confidence in the SR findings. To examine the challenges of using LLMs for QS, we conducted a collaborative autoethnography involving two trials. We evaluated each trial for methodological rigor and practical usefulness, and interpreted the results through a technical lens informed by how LLMs are built and their current limitations.
