Readability Reconsidered: A Cross-Dataset Analysis of Reference-Free Metrics
Catarina G Belem, Parker Glenn, Alfy Samuel, Anoop Kumar, Daben Liu
TL;DR
The paper tackles the lack of a unified notion of readability by evaluating a broad set of metrics against human judgments across five English datasets. It interrogates surface-based, psycholinguistic, and model-based approaches, including fine-tuned models, ReadMe++, and LLM-based judges, highlighting that model-based metrics offer superior alignment with human judgments (up to about 0.73 Kendall Tau-b) though at higher inference cost. The findings reveal that information content and topic influence readability, with human rationales showing that factors beyond surface features drive comprehension. The work underscores the need for clearer readability definitions and stronger validation, and it points toward model-based approaches as a promising direction for more accurate, domain-relevant readability assessment, potentially extending to multilingual contexts.
Abstract
Automatic readability assessment plays a key role in ensuring effective and accessible written communication. Despite significant progress, the field is hindered by inconsistent definitions of readability and measurements that rely on surface-level text properties. In this work, we investigate the factors shaping human perceptions of readability through the analysis of 897 judgments, finding that, beyond surface-level cues, information content and topic strongly shape text comprehensibility. Furthermore, we evaluate 15 popular readability metrics across five English datasets, contrasting them with six more nuanced, model-based metrics. Our results show that four model-based metrics consistently place among the top four in rank correlations with human judgments, while the best performing traditional metric achieves an average rank of 8.6. These findings highlight a mismatch between current readability metrics and human perceptions, pointing to model-based approaches as a more promising direction.
