The dynamics of discovery and the Heaps-Zipf relationship
Célestin Zimmerlin, Thomas Louail, Manuel Moussallam, Marc Barthelemy
TL;DR
The paper questions the universality of the Zipf–Heaps relationship and investigates how temporal correlations sculpt type–token growth in real sequences. By contrasting ordered sequences with reshuffled copies across text, music listening, and web data, and by fitting $D = c k^\alpha$ to obtain $\alpha$ and $\alpha^*$, the authors show that temporal structure can decouple the type–token curve from the rank–frequency distribution $p(r)$, especially in music and web scenarios. They introduce a minimal one-parameter toy model with memory that reproduces a wide range of $D(k)$ trajectories and define an envelope of all admissible Heaps curves for a fixed $p(r)$. The findings urge caution in interpreting $\alpha$ as a universal discovery rate and highlight the need to account for domain-specific temporal dynamics when applying Zipf–Heaps analyses to human behavior.
Abstract
When following a sequence - such as reading a text or tracking a user's activity - one can measure how the "dictionary" of distinct elements (types) grows with the number of observations (tokens). When this growth follows a power law, it is referred to as Heaps' law, a regularity often associated with Zipf's law and frequently used to characterize human innovation and discovery processes. While random sampling from a Zipf-like distribution can reproduce Heaps' law, this connection relies on the assumption of temporal independence - an assumption often violated in real-world systems although frequently found in the literature. Here, we investigate how temporal correlations in token sequences affect the type-token curve. In systems like music listening and web browsing, domain-specific correlations in token ordering lead to systematic deviations from the Zipf-Heaps framework, effectively decoupling the type-token plot from the rank-frequency distribution. Using a minimal one-parameter model, we reproduce a wide variety of type-token trajectories, including the extremal cases that bound all possible behaviors compatible with a given frequency distribution. Our results demonstrate that type-token growth reflects not only the empirical distribution of type frequencies, but also the temporal structure of the sequence - a factor often overlooked in empirical applications of scaling laws to characterize human behavior.
