Table of Contents
Fetching ...

Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers

Tuhin Chakrabarty, Jane C. Ginsburg, Paramveer Dhillon

TL;DR

The study investigates whether AI trained on copyrighted books can rival expert human writers in literary quality and stylistic fidelity. In-context prompting initially yielded human-preferred writing among experts, but author-specific fine-tuning flipped this, with AI outputs outperforming human authors for both style and quality across reader groups and author styles. Detectability of AI-generated text declined dramatically after fine-tuning, and economic analysis showed substantial cost savings (median ≈$81 per author) for producing publishable prose compared with hiring professional writers. The findings imply that fine-tuned AI emulations can substantially disrupt literary markets and inform fair-use considerations, suggesting governance measures and guardrails are needed to manage market dilution while preserving legitimate uses of AI-generated text.

Abstract

The use of copyrighted books for training AI models has led to numerous lawsuits from authors concerned about AI's ability to generate derivative content. Yet it's unclear if these models can generate high quality literary text while emulating authors' styles. To answer this we conducted a preregistered study comparing MFA-trained expert writers with three frontier AI models: ChatGPT, Claude & Gemini in writing up to 450 word excerpts emulating 50 award-winning authors' diverse styles. In blind pairwise evaluations by 159 representative expert & lay readers, AI-generated text from in-context prompting was strongly disfavored by experts for both stylistic fidelity (OR=0.16, p<10^-8) & writing quality (OR=0.13, p<10^-7) but showed mixed results with lay readers. However, fine-tuning ChatGPT on individual authors' complete works completely reversed these findings: experts now favored AI-generated text for stylistic fidelity (OR=8.16, p<10^-13) & writing quality (OR=1.87, p=0.010), with lay readers showing similar shifts. These effects generalize across authors & styles. The fine-tuned outputs were rarely flagged as AI-generated (3% rate v. 97% for in-context prompting) by best AI detectors. Mediation analysis shows this reversal occurs because fine-tuning eliminates detectable AI stylistic quirks (e.g., cliche density) that penalize in-context outputs. While we do not account for additional costs of human effort required to transform raw AI output into cohesive, publishable prose, the median fine-tuning & inference cost of $81 per author represents a dramatic 99.7% reduction compared to typical professional writer compensation. Author-specific fine-tuning thus enables non-verbatim AI writing that readers prefer to expert human writing, providing empirical evidence directly relevant to copyright's fourth fair-use factor, the "effect upon the potential market or value" of the source works.

Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers

TL;DR

The study investigates whether AI trained on copyrighted books can rival expert human writers in literary quality and stylistic fidelity. In-context prompting initially yielded human-preferred writing among experts, but author-specific fine-tuning flipped this, with AI outputs outperforming human authors for both style and quality across reader groups and author styles. Detectability of AI-generated text declined dramatically after fine-tuning, and economic analysis showed substantial cost savings (median ≈$81 per author) for producing publishable prose compared with hiring professional writers. The findings imply that fine-tuned AI emulations can substantially disrupt literary markets and inform fair-use considerations, suggesting governance measures and guardrails are needed to manage market dilution while preserving legitimate uses of AI-generated text.

Abstract

The use of copyrighted books for training AI models has led to numerous lawsuits from authors concerned about AI's ability to generate derivative content. Yet it's unclear if these models can generate high quality literary text while emulating authors' styles. To answer this we conducted a preregistered study comparing MFA-trained expert writers with three frontier AI models: ChatGPT, Claude & Gemini in writing up to 450 word excerpts emulating 50 award-winning authors' diverse styles. In blind pairwise evaluations by 159 representative expert & lay readers, AI-generated text from in-context prompting was strongly disfavored by experts for both stylistic fidelity (OR=0.16, p<10^-8) & writing quality (OR=0.13, p<10^-7) but showed mixed results with lay readers. However, fine-tuning ChatGPT on individual authors' complete works completely reversed these findings: experts now favored AI-generated text for stylistic fidelity (OR=8.16, p<10^-13) & writing quality (OR=1.87, p=0.010), with lay readers showing similar shifts. These effects generalize across authors & styles. The fine-tuned outputs were rarely flagged as AI-generated (3% rate v. 97% for in-context prompting) by best AI detectors. Mediation analysis shows this reversal occurs because fine-tuning eliminates detectable AI stylistic quirks (e.g., cliche density) that penalize in-context outputs. While we do not account for additional costs of human effort required to transform raw AI output into cohesive, publishable prose, the median fine-tuning & inference cost of $81 per author represents a dramatic 99.7% reduction compared to typical professional writer compensation. Author-specific fine-tuning thus enables non-verbatim AI writing that readers prefer to expert human writing, providing empirical evidence directly relevant to copyright's fourth fair-use factor, the "effect upon the potential market or value" of the source works.
Paper Structure (13 sections, 9 equations, 24 figures, 15 tables)

This paper contains 13 sections, 9 equations, 24 figures, 15 tables.

Figures (24)

  • Figure 1: Figure showing our study design. (1) Select a target author and prompt. (2) Generate upto 450-word candidate excerpts from MFA experts and from LLMs under two settings: in-context prompting (instructions + few-shot examples) and author-specific fine-tuning (model fine-tuned on that author’s works). (3) Readers (experts and lay ) perform blinded, pairwise forced-choice evaluations on two outcomes: stylistic fidelity to the target author and overall writing quality. Pair order and left/right placement are randomized on every trial.
  • Figure 2: (A-B) Forest plots showing odds ratios (OR) and 95% confidence intervals comparing AI and human experts in pairwise evaluations of stylistic fidelity (A) and writing quality (B) where values $>$1 favor AI and values $<$1 favor humans. Expert readers show preference for human writing when prompted in an in-context setting (OR = 0.16 and 0.13) but that changes when AI is fine-tuned (OR = 8.16 and 1.87). Lay readers have a harder time discriminating, given how they prefer the quality of AI writing even for in-context prompting (OR = 1.55). (C-D) Probability of choosing AI excerpts across individual language models for stylistic fidelity (C) and writing quality (D). Error bars represent 95% confidence intervals. Dashed line indicates chance performance (50%). (E) AI detection accuracy with chosen threshold of $\tau$=0.9 using two state-of-the-art AI detectors (Pangram and GPTZero). Human written text was never misclassified (0.00), in-context AI was detected with 97% accuracy by Pangram and 91% by GPTZero, but fine-tuned AI evaded detection 97% of the time (0.03) for Pangram and 100% of the time in case of GPTZero. (F) Relationship between AI detectability (Pangram) and preference for writing quality across detection score bins. For in-context prompting setup, higher detection scores correlated with lower AI preference (negative slope). This relationship disappeared when AI is fine-tuned (flat slopes). Fine-tuning on an author's complete oeuvre gets rid of AI quirks while achieving expert-level performance. n = 28 expert readers, 131 lay readers; 3,840 pairwise comparisons with robust clustered standard errors.
  • Figure 3: Author-level AI preference and its association with fine-tuning corpus size (fidelity and quality) (A) For each fine-tuned author, the share of blinded pairwise trials in which the AI excerpt was preferred over the human (MFA expert) on stylistic fidelity (top) and overall quality (bottom). Points show Jeffreys-prior estimates $(k+0.5)/(n+1)$; vertical bars are 95% Jeffreys intervals (Beta$(\tfrac{1}{2},\tfrac{1}{2})$); the dotted line at 0.5 marks human–AI parity. Readers are pooled (experts and lay). (B) AI preference rate versus the fine-tuning corpus size for that author (million tokens), shown for stylistic fidelity (top) and overall quality (bottom). Each point is a fine-tuned author; the line is an OLS fit with heteroskedasticity-robust standard errors (no CI displayed). Slopes are near zero in both panels, indicating little association between corpus size (in this range) and AI preference.
  • Figure 4: Fine-tuning substantially reduces stylometric signatures of AI text, improves stylistic fidelity and perceived writing quality over in-context prompting, and substantially cuts costs of producing first draft versus professional writers. (A) Mediation analysis linking stylometric features $\rightarrow$ AI-detector score $\rightarrow$ human preference. Standardized logistic coefficients with 95% CIs are shown for three features for in-context prompting (red) and fine-tuned models (green). Cliché density mediates 16.4% of the detector effect on choice for in-context prompting but only 1.3% for author fine-tuned models; all three features together mediate 25.4% vs −3.2% for in-context prompted and fine-tuned models respectively. (B) "Fine-tuning premium," defined as $\Delta$ = P(prefer fine-tuned over human) − P(prefer in-context over human), as a function of fine-tuning corpus size. Top: stylistic fidelity; bottom: writing quality. Points are authors; colors denote improvement (green, $\Delta>0$), no change (gray), or degradation (red, $\Delta< 0$). Median $\Delta$: +41.7 (fidelity) and +16.7 percentage points (quality). (C) Cost to produce 100,000 words of raw text vs. publishable prose. Expert writers in our study would earn $25,000 for a 100k-word novel-length manuscript (red). By contrast, AI pipelines can generate 100k words of raw text for $25–$276 depending on fine-tuning corpus size (green bars = fine-tuning; hatched = in-context prompting, $3). This figure reflects direct compute/API costs only, not the additional human steering, chunking, and editing required to turn raw AI text into a cohesive publishable work. Authors ordered by total AI cost.
  • Figure 5: In-Context Writing Prompt used in AI Condition 1.
  • ...and 19 more figures