Table of Contents
Fetching ...

Show Your Title! A Scoping Review on Verbalization in Software Engineering with LLM-Assisted Screening

Gergő Balogh, Dávid Kószó, Homayoun Safarpour Motealegh Mahalegi, László Tóth, Bence Szakács, Áron Búcsú

TL;DR

This study addresses the challenge of understanding software developers’ cognition by leveraging verbalization techniques and conducting a scoping review at the software engineering–psychology interface. It demonstrates a novel LLM-assisted screening pipeline that assesses relevance using only paper titles, validated against human judgments, and maps themes across SE and Psy literatures. The analysis reveals prominent themes around expert cognition, decision-making, and collaboration, with notable asymmetry in methodological borrowing between SE and Psy; underrepresented topics include diversity, implicit human factors, and historical context. The work shows the practical feasibility of AI-assisted interdisciplinary scoping and lays groundwork for a detailed follow-up systematic review and a catalog of verbalization techniques linking psychology theory with SE practice.

Abstract

Understanding how software developers think, make decisions, and behave remains a key challenge in software engineering (SE). Verbalization techniques (methods that capture spoken or written thought processes) offer a lightweight and accessible way to study these cognitive aspects. This paper presents a scoping review of research at the intersection of SE and psychology (PSY), focusing on the use of verbal data. To make large-scale interdisciplinary reviews feasible, we employed a large language model (LLM)-assisted screening pipeline using GPT to assess the relevance of over 9,000 papers based solely on titles. We addressed two questions: what themes emerge from verbalization-related work in SE, and how effective are LLMs in supporting interdisciplinary review processes? We validated GPT's outputs against human reviewers and found high consistency, with a 13\% disagreement rate. Prominent themes mainly were tied to the craft of SE, while more human-centered topics were underrepresented. The data also suggests that SE frequently draws on PSY methods, whereas the reverse is rare.

Show Your Title! A Scoping Review on Verbalization in Software Engineering with LLM-Assisted Screening

TL;DR

This study addresses the challenge of understanding software developers’ cognition by leveraging verbalization techniques and conducting a scoping review at the software engineering–psychology interface. It demonstrates a novel LLM-assisted screening pipeline that assesses relevance using only paper titles, validated against human judgments, and maps themes across SE and Psy literatures. The analysis reveals prominent themes around expert cognition, decision-making, and collaboration, with notable asymmetry in methodological borrowing between SE and Psy; underrepresented topics include diversity, implicit human factors, and historical context. The work shows the practical feasibility of AI-assisted interdisciplinary scoping and lays groundwork for a detailed follow-up systematic review and a catalog of verbalization techniques linking psychology theory with SE practice.

Abstract

Understanding how software developers think, make decisions, and behave remains a key challenge in software engineering (SE). Verbalization techniques (methods that capture spoken or written thought processes) offer a lightweight and accessible way to study these cognitive aspects. This paper presents a scoping review of research at the intersection of SE and psychology (PSY), focusing on the use of verbal data. To make large-scale interdisciplinary reviews feasible, we employed a large language model (LLM)-assisted screening pipeline using GPT to assess the relevance of over 9,000 papers based solely on titles. We addressed two questions: what themes emerge from verbalization-related work in SE, and how effective are LLMs in supporting interdisciplinary review processes? We validated GPT's outputs against human reviewers and found high consistency, with a 13\% disagreement rate. Prominent themes mainly were tied to the craft of SE, while more human-centered topics were underrepresented. The data also suggests that SE frequently draws on PSY methods, whereas the reverse is rare.
Paper Structure (15 sections, 3 figures)

This paper contains 15 sections, 3 figures.

Figures (3)

  • Figure 1: Flowchart of our ScR process with the five stages
  • Figure 2: Normalized distribution of the inclusion question tags of the relevant papers
  • Figure 3: Flowchart of our ScR's validation process (grey = review process step, bold black = validation step)