Symmetry and Generalisation in Neural Approximations of Renormalisation Transformations
Cassidy Ashworth, Pietro Liò, Francesco Caso
TL;DR
The paper investigates how symmetry constraints in neural networks interact with their expressivity when learning renormalisation-group–like decimation maps, using the central limit theorem as a controlled test case. It develops a cumulant-propagation framework to analyse how distributions transform through MLPs and extends this to GNNs, revealing a tension: overly symmetric or overly expressive models can generalise poorly. analytically shows that symmetrically constrained weights can overconstrain simple MLPs with nonlinear activations (e.g., quadratic) and hinder learning, while allowing asymmetry or limited nonlinearity can improve generalisation on CLT-inspired tasks. Empirically, the study demonstrates that the task structure dictates whether symmetry constraints help or hurt, with spline and GNN biases often acting as Occam’s razor toward minimal nonlinearity. The findings offer guidance for encoding physical priors in neural models and suggest adaptive symmetry approaches for robust learning in physics-inspired tasks.
Abstract
Deep learning models have proven enormously successful at using multiple layers of representation to learn relevant features of structured data. Encoding physical symmetries into these models can improve performance on difficult tasks, and recent work has motivated the principle of parameter symmetry breaking and restoration as a unifying mechanism underlying their hierarchical learning dynamics. We evaluate the role of parameter symmetry and network expressivity in the generalisation behaviour of neural networks when learning a real-space renormalisation group (RG) transformation, using the central limit theorem (CLT) as a test case map. We consider simple multilayer perceptrons (MLPs) and graph neural networks (GNNs), and vary weight symmetries and activation functions across architectures. Our results reveal a competition between symmetry constraints and expressivity, with overly complex or overconstrained models generalising poorly. We analytically demonstrate this poor generalisation behaviour for certain constrained MLP architectures by recasting the CLT as a cumulant recursion relation and making use of an established framework to propagate cumulants through MLPs. We also empirically validate an extension of this framework from MLPs to GNNs, elucidating the internal information processing performed by these more complex models. These findings offer new insight into the learning dynamics of symmetric networks and their limitations in modelling structured physical transformations.
