Table of Contents
Fetching ...

Evolution of the lexicon: a probabilistic point of view

Maurizio Serva

TL;DR

This work analyzes the probabilistic limits of traditional Swadesh glottochronology and introduces a complementary stochastic view of lexicon evolution, showing that finite Swadesh lists impose inherent dating uncertainty even under ideal conditions. By coupling word replacement at rate $\\lambda$ with gradual word modification at rate $\\mu$, the authors derive explicit expressions for cognate overlap and normalized distances, enabling three dating strategies: cognate-based, blind Hamming-based, and cognate-aware hybrids. They demonstrate that, for times up to about 3,000 years, a blind approach using normalized edit distance is both more tractable and often more accurate than cognacy-based methods, while longer times favor mixed or cognate-focused strategies albeit with greater practical challenges. The framework provides explicit error bounds and generalizes to broad applications beyond linguistics, including automatic distance-based analyses in diverse domains.

Abstract

The Swadesh approach for determining the temporal separation between two languages relies on the stochastic process of words replacement (when a complete new word emerges to represent a given concept). It is well known that the basic assumptions of the Swadesh approach are often unrealistic due to various contamination phenomena and misjudgments (horizontal transfers, variations over time and space of the replacement rate, incorrect assessments of cognacy relationships, presence of synonyms, and so on). All of this means that the results cannot be completely correct. More importantly, even in the unrealistic case that all basic assumptions are satisfied, simple mathematics places limits on the accuracy of estimating the temporal separation between two languages. These limits, which are purely probabilistic in nature and which are often neglected in lexicostatistical studies, are analyzed in detail in this article. Furthermore, in this work we highlight that the evolution of a language's lexicon is also driven by another stochastic process: gradual lexical modification of words. We show that this process equally also represents a major contribution to the reshaping of the vocabulary of languages over the centuries and we also show, from a purely probabilistic perspective, that taking into account this second random process significantly increases the precision in determining the temporal separation between two languages.

Evolution of the lexicon: a probabilistic point of view

TL;DR

This work analyzes the probabilistic limits of traditional Swadesh glottochronology and introduces a complementary stochastic view of lexicon evolution, showing that finite Swadesh lists impose inherent dating uncertainty even under ideal conditions. By coupling word replacement at rate with gradual word modification at rate , the authors derive explicit expressions for cognate overlap and normalized distances, enabling three dating strategies: cognate-based, blind Hamming-based, and cognate-aware hybrids. They demonstrate that, for times up to about 3,000 years, a blind approach using normalized edit distance is both more tractable and often more accurate than cognacy-based methods, while longer times favor mixed or cognate-focused strategies albeit with greater practical challenges. The framework provides explicit error bounds and generalizes to broad applications beyond linguistics, including automatic distance-based analyses in diverse domains.

Abstract

The Swadesh approach for determining the temporal separation between two languages relies on the stochastic process of words replacement (when a complete new word emerges to represent a given concept). It is well known that the basic assumptions of the Swadesh approach are often unrealistic due to various contamination phenomena and misjudgments (horizontal transfers, variations over time and space of the replacement rate, incorrect assessments of cognacy relationships, presence of synonyms, and so on). All of this means that the results cannot be completely correct. More importantly, even in the unrealistic case that all basic assumptions are satisfied, simple mathematics places limits on the accuracy of estimating the temporal separation between two languages. These limits, which are purely probabilistic in nature and which are often neglected in lexicostatistical studies, are analyzed in detail in this article. Furthermore, in this work we highlight that the evolution of a language's lexicon is also driven by another stochastic process: gradual lexical modification of words. We show that this process equally also represents a major contribution to the reshaping of the vocabulary of languages over the centuries and we also show, from a purely probabilistic perspective, that taking into account this second random process significantly increases the precision in determining the temporal separation between two languages.
Paper Structure (6 sections, 71 equations, 2 figures)

This paper contains 6 sections, 71 equations, 2 figures.

Figures (2)

  • Figure 1: The relative errors $R_\omega$ (red) as defined by (\ref{['sub7']}), concerning the classical Swadesh approach, $R_\phi$ (green) as defined by (\ref{['leb1']}), (\ref{['leb6']}) and (\ref{['A15']}), concerning the blind use of normalized edit distance and $R_\varphi$ (blue) as defined by (\ref{['me3']}), (\ref{['me6']}) and (\ref{['A11']}), concerning the combined use of edit distance and cognate identification. We plot the three errors in function of $T$ in the interval $300 \le T \le 6,000$. We use the Swadesh proposal for the replacement rate $\lambda =1.4 \times10^{-4}$ which is confirmed by our estimate. We assume $M=207$, again according to Swadesh. Moreover, we use our estimates $\mu=1.6 \times10^{-4}$, $N=5.18$ and $L=7.63$.
  • Figure 2: The parameters $\lambda(g)$ (blue), $\hat{\mu}(g) = \frac{N\!\!-\!\!1}{N} \mu(g)$ (green) and $R(g)$ (red) as a function of the minimal geographical distance $g$ (in $Km$) between two varieties in a pair. The parameter $\lambda(g)$ is the rate of words replacements, $\hat{\mu}(g)$ is the rate of effective characters changes and $R(g)$ is the number of language pairs involved in the average, $i.e.,$ the number of language pairs which match at the root of the family tree and whose geographical distance is larger than $g$.