Degree distributions in networks: beyond the power law
Clement Lee, Emma Eastoe, Aiden Farrell
TL;DR
The paper tackles the limitations of using a single power-law to describe network degree distributions, notably threshold selection and adequacy testing. It proposes a Bayesian extreme-value mixture that combines a discrete TZP body with a discrete GP tail via an IGP, allowing threshold uncertainty to be quantified and a formal test for power-law adequacy. By employing 2- and 3-component TZP-IGP mixtures and a spike-and-slab-based model selection, the approach reveals when the body follows a power law and how tails deviate, across real-world networks and word-frequency data. The results show strong goodness-of-fit for many datasets, with the mixture capturing piecewise linear survival curves and correcting tail behavior that Zipf-based models miss, offering a principled alternative to preferential attachment as a network-generating mechanism.
Abstract
The power law is useful in describing count phenomena such as network degrees and word frequencies. With a single parameter, it captures the main feature that the frequencies are linear on the log-log scale. Nevertheless, there have been criticisms of the power law, for example that a threshold needs to be pre-selected without its uncertainty quantified, that the power law is simply inadequate, and that subsequent hypothesis tests are required to determine whether the data could have come from the power law. We propose a modelling framework that combines two different generalisations of the power law, namely the generalised Pareto distribution and the Zipf-polylog distribution, to resolve these issues. The proposed mixture distributions are shown to fit the data well and quantify the threshold uncertainty in a natural way. A model selection step embedded in the Bayesian inference algorithm further answers the question whether the power law is adequate.
