Memory capacity of two layer neural networks with smooth activations

Liam Madden; Christos Thrampoulidis

Memory capacity of two layer neural networks with smooth activations

Liam Madden, Christos Thrampoulidis

TL;DR

This work analyzes the memory capacity of two-layer neural networks with $m$ hidden units and input dimension $d$, showing a lower bound of $\lfloor md/2\rfloor$ on the data size that can be memorized when activations are real analytic at a point, and proving near-optimality up to a factor of $\approx 2$. The authors develop a general framework based on the Jacobian with respect to the first-layer weights and leverage a decomposition into Hadamard powers and the Khatri-Rao product to obtain precise generic-rank results for Hadamard powers, their polynomial sums, and their Khatri-Rao combinations, with extensions to real-analytic Hadamard functions. These results yield concrete surjectivity and non-surjectivity conditions (e.g., surjectivity when $md\ge 2n$ and $m$ even; non-surjectivity when $m(d+2)<n$) and provide a principled interpolation mechanism that applies to deeper models via a constant-rank argument. By linking the algebraic-rank structure to memory capacity, the paper advances a rigorous understanding of interpolation thresholds for smooth activations and lays groundwork for extending the approach to deeper architectures and other linkages.

Abstract

Determining the memory capacity of two layer neural networks with $m$ hidden neurons and input dimension $d$ (i.e., $md+2m$ total trainable parameters), which refers to the largest size of general data the network can memorize, is a fundamental machine learning question. For activations that are real analytic at a point and, if restricting to a polynomial there, have sufficiently high degree, we establish a lower bound of $\lfloor md/2\rfloor$ and optimality up to a factor of approximately $2$. All practical activations, such as sigmoids, Heaviside, and the rectified linear unit (ReLU), are real analytic at a point. Furthermore, the degree condition is mild, requiring, for example, that $\binom{k+d-1}{d-1}\ge n$ if the activation is $x^k$. Analogous prior results were limited to Heaviside and ReLU activations -- our result covers almost everything else. In order to analyze general activations, we derive the precise generic rank of the network's Jacobian, which can be written in terms of Hadamard powers and the Khatri-Rao product. Our analysis extends classical linear algebraic facts about the rank of Hadamard powers. Overall, our approach differs from prior works on memory capacity and holds promise for extending to deeper models and other architectures.

Memory capacity of two layer neural networks with smooth activations

TL;DR

This work analyzes the memory capacity of two-layer neural networks with

hidden units and input dimension

, showing a lower bound of

on the data size that can be memorized when activations are real analytic at a point, and proving near-optimality up to a factor of

. The authors develop a general framework based on the Jacobian with respect to the first-layer weights and leverage a decomposition into Hadamard powers and the Khatri-Rao product to obtain precise generic-rank results for Hadamard powers, their polynomial sums, and their Khatri-Rao combinations, with extensions to real-analytic Hadamard functions. These results yield concrete surjectivity and non-surjectivity conditions (e.g., surjectivity when

and

even; non-surjectivity when

) and provide a principled interpolation mechanism that applies to deeper models via a constant-rank argument. By linking the algebraic-rank structure to memory capacity, the paper advances a rigorous understanding of interpolation thresholds for smooth activations and lays groundwork for extending the approach to deeper architectures and other linkages.

Abstract

Determining the memory capacity of two layer neural networks with

hidden neurons and input dimension

(i.e.,

total trainable parameters), which refers to the largest size of general data the network can memorize, is a fundamental machine learning question. For activations that are real analytic at a point and, if restricting to a polynomial there, have sufficiently high degree, we establish a lower bound of

and optimality up to a factor of approximately

. All practical activations, such as sigmoids, Heaviside, and the rectified linear unit (ReLU), are real analytic at a point. Furthermore, the degree condition is mild, requiring, for example, that

if the activation is

. Analogous prior results were limited to Heaviside and ReLU activations -- our result covers almost everything else. In order to analyze general activations, we derive the precise generic rank of the network's Jacobian, which can be written in terms of Hadamard powers and the Khatri-Rao product. Our analysis extends classical linear algebraic facts about the rank of Hadamard powers. Overall, our approach differs from prior works on memory capacity and holds promise for extending to deeper models and other architectures.

Paper Structure (23 sections, 13 theorems, 67 equations)

This paper contains 23 sections, 13 theorems, 67 equations.

Introduction
Results
Related work
Organization
Preliminaries
Rank of a matrix
Real analytic functions of several variables
Weak compositions
Results involving Hadamard powers
Proof sketch
Matrix products
Hadamard powers
Polynomial Hadamard functions
Results involving the Khatri-Rao product
Hadamard powers with Khatri-Rao products
...and 8 more sections

Key Result

Lemma 2.1

Let $A\in\mathbb R^{m\times N}$ and $B\in\mathbb R^{n\times N}$, and let $D\in\mathbb R^{N\times N}$ be diagonal. Let $I\subset [m]$ and $J\subset [n]$ such that $|I|=|J|=s\le \min\{m,n\}$. Then

Theorems & Definitions (23)

Lemma 2.1
proof
Lemma 2.2
Theorem 3.1
proof
Theorem 3.2
proof
Lemma 4.1
proof
Theorem 4.2
...and 13 more

Memory capacity of two layer neural networks with smooth activations

TL;DR

Abstract

Memory capacity of two layer neural networks with smooth activations

Authors

TL;DR

Abstract

Table of Contents

Key Result

Theorems & Definitions (23)