Table of Contents
Fetching ...

Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics

Catarina G Belem, Parker Glenn, Alfy Samuel, Anoop Kumar, Daben Liu

TL;DR

The paper tackles the lack of a unified notion of readability by evaluating a broad set of metrics against human judgments across five English datasets. It interrogates surface-based, psycholinguistic, and model-based approaches, including fine-tuned models, ReadMe++, and LLM-based judges, highlighting that model-based metrics offer superior alignment with human judgments (up to about 0.73 Kendall Tau-b) though at higher inference cost. The findings reveal that information content and topic influence readability, with human rationales showing that factors beyond surface features drive comprehension. The work underscores the need for clearer readability definitions and stronger validation, and it points toward model-based approaches as a promising direction for more accurate, domain-relevant readability assessment, potentially extending to multilingual contexts.

Abstract

Automatic readability assessment plays a key role in ensuring effective and accessible written communication. Despite significant progress, the field is hindered by inconsistent definitions of readability and measurements that rely on surface-level text properties. In this work, we investigate the factors shaping human perceptions of readability through the analysis of 897 judgments, finding that, beyond surface-level cues, information content and topic strongly shape text comprehensibility. Furthermore, we evaluate 15 popular readability metrics across five English datasets, contrasting them with six more nuanced, model-based metrics. Our results show that four model-based metrics consistently place among the top four in rank correlations with human judgments, while the best performing traditional metric achieves an average rank of 8.6. These findings highlight a mismatch between current readability metrics and human perceptions, pointing to model-based approaches as a more promising direction.

Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics

TL;DR

The paper tackles the lack of a unified notion of readability by evaluating a broad set of metrics against human judgments across five English datasets. It interrogates surface-based, psycholinguistic, and model-based approaches, including fine-tuned models, ReadMe++, and LLM-based judges, highlighting that model-based metrics offer superior alignment with human judgments (up to about 0.73 Kendall Tau-b) though at higher inference cost. The findings reveal that information content and topic influence readability, with human rationales showing that factors beyond surface features drive comprehension. The work underscores the need for clearer readability definitions and stronger validation, and it points toward model-based approaches as a promising direction for more accurate, domain-relevant readability assessment, potentially extending to multilingual contexts.

Abstract

Automatic readability assessment plays a key role in ensuring effective and accessible written communication. Despite significant progress, the field is hindered by inconsistent definitions of readability and measurements that rely on surface-level text properties. In this work, we investigate the factors shaping human perceptions of readability through the analysis of 897 judgments, finding that, beyond surface-level cues, information content and topic strongly shape text comprehensibility. Furthermore, we evaluate 15 popular readability metrics across five English datasets, contrasting them with six more nuanced, model-based metrics. Our results show that four model-based metrics consistently place among the top four in rank correlations with human judgments, while the best performing traditional metric achieves an average rank of 8.6. These findings highlight a mismatch between current readability metrics and human perceptions, pointing to model-based approaches as a more promising direction.
Paper Structure (27 sections, 9 equations, 6 figures, 12 tables)

This paper contains 27 sections, 9 equations, 6 figures, 12 tables.

Figures (6)

  • Figure 1: Distribution of justification reasons across 90 examples in ELI-Why (GPT-4). Counts are based on the consensus over 2-way annotations.
  • Figure 2: Formatting of each ScienceQA example. Whenever examples miss the corresponding {{lecture}} or {{explanation}} fields, we we omit them from the template above.
  • Figure 3: Distribution of number of words (# Words) and sentences (# Sentences) per readability label in the ELI-Why (GPT-4) dataset.
  • Figure 4: Prompt used to extract a 0-100 continuous score associated with the ease of readability of a given text. The placeholder {{text}} is either the explanation to a question or the text excerpts depending on the dataset being evaluated.
  • Figure 5: Frequency-based analysis of the language expressions used by human annotators when judging the perceived readability of various GPT4-generated explanations in ELI-Why (GPT-4). These word clouds are collected over 324, 694, and 182 examples annotated for Elementary, High School, and Graduate, respectively.
  • ...and 1 more figures