Table of Contents
Fetching ...

The dynamics of discovery and the Heaps-Zipf relationship

Célestin Zimmerlin, Thomas Louail, Manuel Moussallam, Marc Barthelemy

TL;DR

The paper questions the universality of the Zipf–Heaps relationship and investigates how temporal correlations sculpt type–token growth in real sequences. By contrasting ordered sequences with reshuffled copies across text, music listening, and web data, and by fitting $D = c k^\alpha$ to obtain $\alpha$ and $\alpha^*$, the authors show that temporal structure can decouple the type–token curve from the rank–frequency distribution $p(r)$, especially in music and web scenarios. They introduce a minimal one-parameter toy model with memory that reproduces a wide range of $D(k)$ trajectories and define an envelope of all admissible Heaps curves for a fixed $p(r)$. The findings urge caution in interpreting $\alpha$ as a universal discovery rate and highlight the need to account for domain-specific temporal dynamics when applying Zipf–Heaps analyses to human behavior.

Abstract

When following a sequence - such as reading a text or tracking a user's activity - one can measure how the "dictionary" of distinct elements (types) grows with the number of observations (tokens). When this growth follows a power law, it is referred to as Heaps' law, a regularity often associated with Zipf's law and frequently used to characterize human innovation and discovery processes. While random sampling from a Zipf-like distribution can reproduce Heaps' law, this connection relies on the assumption of temporal independence - an assumption often violated in real-world systems although frequently found in the literature. Here, we investigate how temporal correlations in token sequences affect the type-token curve. In systems like music listening and web browsing, domain-specific correlations in token ordering lead to systematic deviations from the Zipf-Heaps framework, effectively decoupling the type-token plot from the rank-frequency distribution. Using a minimal one-parameter model, we reproduce a wide variety of type-token trajectories, including the extremal cases that bound all possible behaviors compatible with a given frequency distribution. Our results demonstrate that type-token growth reflects not only the empirical distribution of type frequencies, but also the temporal structure of the sequence - a factor often overlooked in empirical applications of scaling laws to characterize human behavior.

The dynamics of discovery and the Heaps-Zipf relationship

TL;DR

The paper questions the universality of the Zipf–Heaps relationship and investigates how temporal correlations sculpt type–token growth in real sequences. By contrasting ordered sequences with reshuffled copies across text, music listening, and web data, and by fitting to obtain and , the authors show that temporal structure can decouple the type–token curve from the rank–frequency distribution , especially in music and web scenarios. They introduce a minimal one-parameter toy model with memory that reproduces a wide range of trajectories and define an envelope of all admissible Heaps curves for a fixed . The findings urge caution in interpreting as a universal discovery rate and highlight the need to account for domain-specific temporal dynamics when applying Zipf–Heaps analyses to human behavior.

Abstract

When following a sequence - such as reading a text or tracking a user's activity - one can measure how the "dictionary" of distinct elements (types) grows with the number of observations (tokens). When this growth follows a power law, it is referred to as Heaps' law, a regularity often associated with Zipf's law and frequently used to characterize human innovation and discovery processes. While random sampling from a Zipf-like distribution can reproduce Heaps' law, this connection relies on the assumption of temporal independence - an assumption often violated in real-world systems although frequently found in the literature. Here, we investigate how temporal correlations in token sequences affect the type-token curve. In systems like music listening and web browsing, domain-specific correlations in token ordering lead to systematic deviations from the Zipf-Heaps framework, effectively decoupling the type-token plot from the rank-frequency distribution. Using a minimal one-parameter model, we reproduce a wide variety of type-token trajectories, including the extremal cases that bound all possible behaviors compatible with a given frequency distribution. Our results demonstrate that type-token growth reflects not only the empirical distribution of type frequencies, but also the temporal structure of the sequence - a factor often overlooked in empirical applications of scaling laws to characterize human behavior.
Paper Structure (10 sections, 5 equations, 4 figures, 1 table)

This paper contains 10 sections, 5 equations, 4 figures, 1 table.

Figures (4)

  • Figure 1: Top: Type–token plots for a random sample from each dataset: texts from the Project Gutenberg corpus (left), individual music listening histories of Deezer users (middle),and individual web browsing histories (right). For each case, both the empirical ordered trajectory and a reshuffled trajectory are shown. The scaling relation $D = c k^\alpha$ (\ref{['eq:Heaps']}) is fitted using two alternative methods; for each sequence, the exponent corresponding to the fit with the lowest Bayesian Information Criterion (BIC) is retained (see Appendix B). Top: Distributions of scaling exponents $\alpha$ obtained for both ordered and reshuffled sequences $D(k)$, across the three datasets. We consider more than 1000 sequences for both the Gutenberg and Deezer datasets, and 292 sequences for the Web Tracking dataset. Error bars are computed using the bootstrap method described in sohil_introduction_2022, with 1,000 bootstrap samples. Bottom: Relationship between exponent values estimated for the ordered ($\alpha$) and reshuffled ($\alpha^{*}$) sequences of the same individual user or text. Each plot reports the Bravais–Pearson correlation coefficient $r$, along with a linear regression $\alpha^{*} = a \alpha + b$ fitted by minimizing the absolute deviation to reduce sensitivity to outliers.
  • Figure 2: Evolution of the absolute difference $|\alpha - \alpha^*|$ as a function of the sequence length $k_{\max}$. Each sequence (either a text or an individual listening history) is truncated at 20 different values of $k_{\max}$, ranging from 10,000 to 40,000 tokens. For each truncation point, we compute the absolute difference between the exponent obtained for the ordered sequence ($\alpha$) and the one obtained for the reshuffled sequence ($\alpha^*$). The resulting values of $|\alpha - \alpha^*|$ are displayed as a heatmap. To highlight the overall trend, we further compute, for each $k_{\max}$, the average $|\alpha - \alpha^*|$ across all sequences and fit a linear regression to these averages. The analysis is based on 500 texts and 200 individual listening histories. The web browsing dataset is excluded due to insufficient sequence length (see Appendix A).
  • Figure 3: Top: Type–token trajectories for different temporal arrangements of a sequence sampled from $p(r) \propto r^{-1.5}$, showing variability in discovery dynamics for the same underlying distribution. The black and red dashed lines correspond to the limiting cases of maximally accelerated and delayed discoveries. Bottom: Corresponding rank–frequency distribution of types in the sequence, plotted in log–log scale.
  • Figure 4: Distribution of the coefficient of determination ($R^2$) when fitting the power-law $D=ck^\alpha$ on the type-token plot (\ref{['eq:Heaps']}), for both ordered and reshuffled trajectories across the three datasets. Error bars are computed using a bootstrap procedure with $1000$ resamples, following the approach described in sohil_introduction_2022.