Table of Contents
Fetching ...

The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis

Hoang Pham, The-Anh Ta, Tom Jacobs, Rebekka Burkholz, Long Tran-Thanh

TL;DR

This work addresses the challenge of trainability in sparse neural networks produced by pruning by introducing the Graphon Limit Hypothesis, which posits that pruning masks converge to layer-wise graphons in the infinite-width limit. Building on graph limit theory, the authors define the Graphon NTK, a kernel that encodes connectivity-induced non-uniform learning dynamics, and they show how a constant graphon recovers a scaled standard NTK for Random Pruning. They provide empirical evidence linking Graphon NTK spectra to early training dynamics across several pruning-at-initialisation methods, suggesting that spectral properties of the graphon-structured NTK predict convergence behavior. The framework unifies pruning analysis under graph limits and offers a principled path toward graphon-guided pruning design and deeper theoretical understanding of sparse network trainability.

Abstract

Sparse neural networks promise efficiency, yet training them effectively remains a fundamental challenge. Despite advances in pruning methods that create sparse architectures, understanding why some sparse structures are better trainable than others with the same level of sparsity remains poorly understood. Aiming to develop a systematic approach to this fundamental problem, we propose a novel theoretical framework based on the theory of graph limits, particularly graphons, that characterizes sparse neural networks in the infinite-width regime. Our key insight is that connectivity patterns of sparse neural networks induced by pruning methods converge to specific graphons as networks' width tends to infinity, which encodes implicit structural biases of different pruning methods. We postulate the Graphon Limit Hypothesis and provide empirical evidence to support it. Leveraging this graphon representation, we derive a Graphon Neural Tangent Kernel (Graphon NTK) to study the training dynamics of sparse networks in the infinite width limit. Graphon NTK provides a general framework for the theoretical analysis of sparse networks. We empirically show that the spectral analysis of Graphon NTK correlates with observed training dynamics of sparse networks, explaining the varying convergence behaviours of different pruning methods. Our framework provides theoretical insights into the impact of connectivity patterns on the trainability of various sparse network architectures.

The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis

TL;DR

This work addresses the challenge of trainability in sparse neural networks produced by pruning by introducing the Graphon Limit Hypothesis, which posits that pruning masks converge to layer-wise graphons in the infinite-width limit. Building on graph limit theory, the authors define the Graphon NTK, a kernel that encodes connectivity-induced non-uniform learning dynamics, and they show how a constant graphon recovers a scaled standard NTK for Random Pruning. They provide empirical evidence linking Graphon NTK spectra to early training dynamics across several pruning-at-initialisation methods, suggesting that spectral properties of the graphon-structured NTK predict convergence behavior. The framework unifies pruning analysis under graph limits and offers a principled path toward graphon-guided pruning design and deeper theoretical understanding of sparse network trainability.

Abstract

Sparse neural networks promise efficiency, yet training them effectively remains a fundamental challenge. Despite advances in pruning methods that create sparse architectures, understanding why some sparse structures are better trainable than others with the same level of sparsity remains poorly understood. Aiming to develop a systematic approach to this fundamental problem, we propose a novel theoretical framework based on the theory of graph limits, particularly graphons, that characterizes sparse neural networks in the infinite-width regime. Our key insight is that connectivity patterns of sparse neural networks induced by pruning methods converge to specific graphons as networks' width tends to infinity, which encodes implicit structural biases of different pruning methods. We postulate the Graphon Limit Hypothesis and provide empirical evidence to support it. Leveraging this graphon representation, we derive a Graphon Neural Tangent Kernel (Graphon NTK) to study the training dynamics of sparse networks in the infinite width limit. Graphon NTK provides a general framework for the theoretical analysis of sparse networks. We empirically show that the spectral analysis of Graphon NTK correlates with observed training dynamics of sparse networks, explaining the varying convergence behaviours of different pruning methods. Our framework provides theoretical insights into the impact of connectivity patterns on the trainability of various sparse network architectures.
Paper Structure (40 sections, 4 theorems, 62 equations, 11 figures)

This paper contains 40 sections, 4 theorems, 62 equations, 11 figures.

Key Result

Proposition 1

For a neural network with layers structured by graphons $\mathcal{W}^{(l)}:[0,1]^2 \rightarrow [0,1]$, Lipschitz nonlinearity $\sigma$, and in the limit as $n_1,...,n_{L} \to \infty$, the pre-activations $z^{(l)}(u_l, x)$ at every hidden layer converge to centred Gaussian processes with covariance $ where $(x \cdot x')_{j}$ represents the input correlation at position $j$, the activation covarianc

Figures (11)

  • Figure 1: Graph limit of subnetworks' mask produced by PaI methods at 80% sparsity and the corresponding convergence of graphons via Euclidean distances.
  • Figure 2: The training loss in the first 200 gradient update steps of training sparse networks produced by Random, SNIP, and Synflow with different sparsity levels. At the beginning of the training phase, subnetworks generated by SNIP and Synflow show a faster convergence speed than Random across sparsity levels.
  • Figure 3: Spectral metrics of the Graphon NTK with different graphon functions and sparsity levels.
  • Figure 4: Histogram convergence via Euclidean distance.
  • Figure 5: Graph limit of subnetworks’ mask produced by PaI methods at different sparsity levels in 4-layer networks.
  • ...and 6 more figures

Theorems & Definitions (6)

  • Proposition 1
  • Remark 1
  • Theorem 1: Graphon NTK
  • Remark 2
  • Proposition 2
  • Theorem 2: Graphon NTK