Table of Contents
Fetching ...

On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?

Mingmeng Geng, Thierry Poibeau

TL;DR

This work interrogates the core assumption behind LLM-generated text detectors: what exactly constitutes LLM-generated text? It argues that the detection target is not well-defined and is highly contingent on prompts, model diversity, and human edits, which undermines universal benchmarking. Through a synthesis of definitions, backgrounds, evaluation challenges, and a case-study, the paper demonstrates that detector outputs are highly sensitive to context and that detectors should be treated as references under specific assumptions rather than definitive judgments. It also surveys attack vectors and watermarking, highlighting an ongoing arms race between detection, evasion, and ethical considerations. The study advocates for transparency, AI-literacy, and human-in-the-loop strategies to responsibly deploy detection tools in practice.

Abstract

With the widespread use of large language models (LLMs), many researchers have turned their attention to detecting text generated by them. However, there is no consistent or precise definition of their target, namely "LLM-generated text". Differences in usage scenarios and the diversity of LLMs further increase the difficulty of detection. What is commonly regarded as the detecting target usually represents only a subset of the text that LLMs can potentially produce. Human edits to LLM outputs, together with the subtle influences that LLMs exert on their users, are blurring the line between LLM-generated and human-written text. Existing benchmarks and evaluation approaches do not adequately address the various conditions in real-world detector applications. Hence, the numerical results of detectors are often misunderstood, and their significance is diminishing. Therefore, detectors remain useful under specific conditions, but their results should be interpreted only as references rather than decisive indicators.

On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?

TL;DR

This work interrogates the core assumption behind LLM-generated text detectors: what exactly constitutes LLM-generated text? It argues that the detection target is not well-defined and is highly contingent on prompts, model diversity, and human edits, which undermines universal benchmarking. Through a synthesis of definitions, backgrounds, evaluation challenges, and a case-study, the paper demonstrates that detector outputs are highly sensitive to context and that detectors should be treated as references under specific assumptions rather than definitive judgments. It also surveys attack vectors and watermarking, highlighting an ongoing arms race between detection, evasion, and ethical considerations. The study advocates for transparency, AI-literacy, and human-in-the-loop strategies to responsibly deploy detection tools in practice.

Abstract

With the widespread use of large language models (LLMs), many researchers have turned their attention to detecting text generated by them. However, there is no consistent or precise definition of their target, namely "LLM-generated text". Differences in usage scenarios and the diversity of LLMs further increase the difficulty of detection. What is commonly regarded as the detecting target usually represents only a subset of the text that LLMs can potentially produce. Human edits to LLM outputs, together with the subtle influences that LLMs exert on their users, are blurring the line between LLM-generated and human-written text. Existing benchmarks and evaluation approaches do not adequately address the various conditions in real-world detector applications. Hence, the numerical results of detectors are often misunderstood, and their significance is diminishing. Therefore, detectors remain useful under specific conditions, but their results should be interpreted only as references rather than decisive indicators.
Paper Structure (11 sections, 2 tables)