Table of Contents
Fetching ...

LLM-REVal: Can We Trust LLM Reviewers Yet?

Rui Li, Jia-Chen Gu, Po-Nien Kung, Heming Xia, Junfeng liu, Xiangwen Kong, Zhifang Sui, Nanyun Peng

TL;DR

This study analyzes the fairness risks when large language models (LLMs) act as reviewers within a simulated academic workflow, revealing systematic biases that inflate scores for LLM-authored work and undervalue human-authored work with critical content. Using a multi-round Research-Agent/Review-Agent framework (LLM-REVal), the authors demonstrate patterns such as LLM-authored superiority, revision-driven score boosts, and inevitable rejection of some human submissions, with biases traced to linguistic features and framing of critical statements. Human annotations reveal misalignment between LLM judgments and human judgments, underscoring potential fairness and equity concerns in deploying LLM-based reviewers at scale. Despite these biases, the revision process guided by LLM feedback can improve paper quality for both LLM- and human-authored papers, suggesting a cautious, role-limited use of LLMs to assist early-stage researchers while safeguarding the integrity of the review process.

Abstract

The rapid advancement of large language models (LLMs) has inspired researchers to integrate them extensively into the academic workflow, potentially reshaping how research is practiced and reviewed. While previous studies highlight the potential of LLMs in supporting research and peer review, their dual roles in the academic workflow and the complex interplay between research and review bring new risks that remain largely underexplored. In this study, we focus on how the deep integration of LLMs into both peer-review and research processes may influence scholarly fairness, examining the potential risks of using LLMs as reviewers by simulation. This simulation incorporates a research agent, which generates papers and revises, alongside a review agent, which assesses the submissions. Based on the simulation results, we conduct human annotations and identify pronounced misalignment between LLM-based reviews and human judgments: (1) LLM reviewers systematically inflate scores for LLM-authored papers, assigning them markedly higher scores than human-authored ones; (2) LLM reviewers persistently underrate human-authored papers with critical statements (e.g., risk, fairness), even after multiple revisions. Our analysis reveals that these stem from two primary biases in LLM reviewers: a linguistic feature bias favoring LLM-generated writing styles, and an aversion toward critical statements. These results highlight the risks and equity concerns posed to human authors and academic research if LLMs are deployed in the peer review cycle without adequate caution. On the other hand, revisions guided by LLM reviews yield quality gains in both LLM-based and human evaluations, illustrating the potential of the LLMs-as-reviewers for early-stage researchers and enhancing low-quality papers.

LLM-REVal: Can We Trust LLM Reviewers Yet?

TL;DR

This study analyzes the fairness risks when large language models (LLMs) act as reviewers within a simulated academic workflow, revealing systematic biases that inflate scores for LLM-authored work and undervalue human-authored work with critical content. Using a multi-round Research-Agent/Review-Agent framework (LLM-REVal), the authors demonstrate patterns such as LLM-authored superiority, revision-driven score boosts, and inevitable rejection of some human submissions, with biases traced to linguistic features and framing of critical statements. Human annotations reveal misalignment between LLM judgments and human judgments, underscoring potential fairness and equity concerns in deploying LLM-based reviewers at scale. Despite these biases, the revision process guided by LLM feedback can improve paper quality for both LLM- and human-authored papers, suggesting a cautious, role-limited use of LLMs to assist early-stage researchers while safeguarding the integrity of the review process.

Abstract

The rapid advancement of large language models (LLMs) has inspired researchers to integrate them extensively into the academic workflow, potentially reshaping how research is practiced and reviewed. While previous studies highlight the potential of LLMs in supporting research and peer review, their dual roles in the academic workflow and the complex interplay between research and review bring new risks that remain largely underexplored. In this study, we focus on how the deep integration of LLMs into both peer-review and research processes may influence scholarly fairness, examining the potential risks of using LLMs as reviewers by simulation. This simulation incorporates a research agent, which generates papers and revises, alongside a review agent, which assesses the submissions. Based on the simulation results, we conduct human annotations and identify pronounced misalignment between LLM-based reviews and human judgments: (1) LLM reviewers systematically inflate scores for LLM-authored papers, assigning them markedly higher scores than human-authored ones; (2) LLM reviewers persistently underrate human-authored papers with critical statements (e.g., risk, fairness), even after multiple revisions. Our analysis reveals that these stem from two primary biases in LLM reviewers: a linguistic feature bias favoring LLM-generated writing styles, and an aversion toward critical statements. These results highlight the risks and equity concerns posed to human authors and academic research if LLMs are deployed in the peer review cycle without adequate caution. On the other hand, revisions guided by LLM reviews yield quality gains in both LLM-based and human evaluations, illustrating the potential of the LLMs-as-reviewers for early-stage researchers and enhancing low-quality papers.
Paper Structure (51 sections, 10 figures, 8 tables)

This paper contains 51 sections, 10 figures, 8 tables.

Figures (10)

  • Figure 1: Pipeline and composition of our simulation. In the research-review round, we have human-author papers and LLM-authored papers generated by the research agent as submissions. The review agent then reviews each paper, and the acceptance decision is made based on the review scores. In the revise-review rounds, we take the low-scoring papers from the previous round, revise them guided by LLM reviews, then repeat the same review process as before.
  • Figure 2: Correlation between LLM review scores and human scores. The box plots illustrate the distribution of LLM review scores across different ranges of human review scores.
  • Figure 2: Average review scores for original (LLM Paper, Human Paper) and their first revisions ($\text{Revision-L}_{1}$, $\text{Revision-H}_{1}$)
  • Figure 3: Review score distributions for papers on various topics. Box plots show the review score distributions for human-authored papers and LLM-authored papers across ten different topics.
  • Figure 4: (Left) Review score distributions for Original Submissions vs First Revision. (Right) Number of submissions in Round 2-6.
  • ...and 5 more figures