Table of Contents
Fetching ...

Statistics of correlations in nonlinear recurrent neural networks

German Mato, Facundo Rigatuso, Gonzalo Torroba

TL;DR

This work develops a path-integral, replica-based framework to compute exact correlation statistics in nonlinear recurrent neural networks in the large-$N$ limit, including $1/N$ corrections. Nonlinear activation functions are incorporated as interaction terms that regularize the linear instability and yield a strictly positive participation dimension, with explicit results for power-law and Padé activations. The theory derives a self-consistent equation for the leading two-point function $G_0$, provides expressions for covariance statistics, and shows how cross-neural correlations scale at finite $N$ while remaining controlled. The approach unifies and extends prior linear analyses, connects to DMFT and random-matrix perspectives, and delivers testable predictions that agree with numerical simulations for networks of a few hundred neurons. It offers a versatile tool for interpreting neural correlations and dimensionality in both neuroscience and machine-learning contexts.

Abstract

The statistics of correlations are central quantities characterizing the collective dynamics of recurrent neural networks. We derive exact expressions for the statistics of correlations of nonlinear recurrent networks in the limit of a large number N of neurons, including systematic 1/N corrections. Our approach uses a path-integral representation of the network's stochastic dynamics, which reduces the description to a few collective variables and enables efficient computation. This generalizes previous results on linear networks to include a wide family of nonlinear activation functions, which enter as interaction terms in the path integral. These interactions can resolve the instability of the linear theory and yield a strictly positive participation dimension. We present explicit results for power-law activations, revealing scaling behavior controlled by the network coupling. In addition, we introduce a class of activation functions based on Pade approximants and provide analytic predictions for their correlation statistics. Numerical simulations confirm our theoretical results with excellent agreement.

Statistics of correlations in nonlinear recurrent neural networks

TL;DR

This work develops a path-integral, replica-based framework to compute exact correlation statistics in nonlinear recurrent neural networks in the large- limit, including corrections. Nonlinear activation functions are incorporated as interaction terms that regularize the linear instability and yield a strictly positive participation dimension, with explicit results for power-law and Padé activations. The theory derives a self-consistent equation for the leading two-point function , provides expressions for covariance statistics, and shows how cross-neural correlations scale at finite while remaining controlled. The approach unifies and extends prior linear analyses, connects to DMFT and random-matrix perspectives, and delivers testable predictions that agree with numerical simulations for networks of a few hundred neurons. It offers a versatile tool for interpreting neural correlations and dimensionality in both neuroscience and machine-learning contexts.

Abstract

The statistics of correlations are central quantities characterizing the collective dynamics of recurrent neural networks. We derive exact expressions for the statistics of correlations of nonlinear recurrent networks in the limit of a large number N of neurons, including systematic 1/N corrections. Our approach uses a path-integral representation of the network's stochastic dynamics, which reduces the description to a few collective variables and enables efficient computation. This generalizes previous results on linear networks to include a wide family of nonlinear activation functions, which enter as interaction terms in the path integral. These interactions can resolve the instability of the linear theory and yield a strictly positive participation dimension. We present explicit results for power-law activations, revealing scaling behavior controlled by the network coupling. In addition, we introduce a class of activation functions based on Pade approximants and provide analytic predictions for their correlation statistics. Numerical simulations confirm our theoretical results with excellent agreement.
Paper Structure (17 sections, 99 equations, 7 figures)

This paper contains 17 sections, 99 equations, 7 figures.

Figures (7)

  • Figure 1: Comparison of analytical (dashed line) and numerical results (dots) for the output correlations. Left panel: average value of the diagonal correlation. Central panel: average of the square of the off-diagonal correlators. Right panel: participation ratio (participation dimension divided by network size). We use the transfer function of Eq.( \ref{['eq:padep0']}) with $\beta=2$ and $D=1$; the network sizes are $N=50, 100, 200$. The error bars represent the standard deviation over 5 simulations with different realizations of the coupling matrix $W$ and noise $\xi$.
  • Figure 2: Scaling with $N$ of the standard deviations of $<C_{ii}^f>$ (left panel) and $N <C_{ij}^{f 2} >$ (right panel). Dashed line shows a fit in powers of $1/N$: $a/N + b/N^2$.
  • Figure 3: Participation ratio for different simulation times. Dashed line: analytical result. Saturating transfer function ($p=0$), $D=1$ and $\beta=2$.
  • Figure 4: Comparison of analytical (dashed line) and numerical results (dots) for the output correlation, for the Padé activation function with $p=1/2$. Left panel: average value of the diagonal correlation. Central panel: average off-diagonal correlators. Right panel: participation ratio (participation dimension divided by network size). We use the transfer function of Eq. \ref{['eq:padep1o2']} with $\beta=2$ and $D=1$. The network sizes are $N=50, 100, 200$. The error bars represent the standard deviation over 5 simulations with different realizations of the coupling matrix $W$ and noise $\xi$.
  • Figure 5: Comparison of participation ratios of inputs and outputs, for $\beta=2, D=1$ (left) and $\beta=2, D=0.1$, right. Note the difference in scales in the y-axis between the plots.
  • ...and 2 more figures