Table of Contents
Fetching ...

Near-Equilibrium Propagation training in nonlinear wave systems

Karol Sajnok, Michał Matuszewski

TL;DR

Stable convergence studies on standard benchmarks, including a simple logical task and handwritten-digit recognition, demonstrate stable convergence, establishing a practical route to in-situ learning in physical systems in which system control is restricted to local parameters.

Abstract

Backpropagation learning algorithm, the workhorse of modern artificial intelligence, is notoriously difficult to implement in physical neural networks. Equilibrium Propagation (EP) is an alternative with comparable efficiency and strong potential for in-situ training. We extend EP learning to both discrete and continuous complex-valued wave systems. In contrast to previous EP implementations, our scheme is valid in the weakly dissipative regime, and readily applicable to a wide range of physical settings, even without well defined nodes, where trainable inter-node connections can be replaced by trainable local potential. We test the method in driven-dissipative exciton-polariton condensates governed by generalized Gross-Pitaevskii dynamics. Numerical studies on standard benchmarks, including a simple logical task and handwritten-digit recognition, demonstrate stable convergence, establishing a practical route to in-situ learning in physical systems in which system control is restricted to local parameters.

Near-Equilibrium Propagation training in nonlinear wave systems

TL;DR

Stable convergence studies on standard benchmarks, including a simple logical task and handwritten-digit recognition, demonstrate stable convergence, establishing a practical route to in-situ learning in physical systems in which system control is restricted to local parameters.

Abstract

Backpropagation learning algorithm, the workhorse of modern artificial intelligence, is notoriously difficult to implement in physical neural networks. Equilibrium Propagation (EP) is an alternative with comparable efficiency and strong potential for in-situ training. We extend EP learning to both discrete and continuous complex-valued wave systems. In contrast to previous EP implementations, our scheme is valid in the weakly dissipative regime, and readily applicable to a wide range of physical settings, even without well defined nodes, where trainable inter-node connections can be replaced by trainable local potential. We test the method in driven-dissipative exciton-polariton condensates governed by generalized Gross-Pitaevskii dynamics. Numerical studies on standard benchmarks, including a simple logical task and handwritten-digit recognition, demonstrate stable convergence, establishing a practical route to in-situ learning in physical systems in which system control is restricted to local parameters.
Paper Structure (26 equations, 7 figures)

This paper contains 26 equations, 7 figures.

Figures (7)

  • Figure 1: Scheme of NEP implementation in a continous wave system (light blue) with input drive, output and nudging force (green spring) and trainable potential parameters (red).
  • Figure 2: Example of a polaritonic system with two inputs and one output. Two localized laser beams (yellow) encode the inputs. A spatially modulated trainable potential (red) is applied via a spatial light modulator. The selected area (green) marks the output region where optical nudging and readout occur. Input and output beams share the same frequency $\omega_D$.
  • Figure 3: Training a 9-node 1D network for the 2-input XOR task. Inputs are applied symmetrically at sites 2 and 6, and the output is read at site 4 (a). Both potential $V$ and pump weights $w$ are trained. (b) Training loss; (c) steady-state field magnitudes $|\Psi_i|^2$ for each input; (e) trained potential $V$; (d) pump weights $w$. Final outputs: $|\Psi_4|^2=(0.00,0.92,1.06,0.01)$ for inputs $(00,01,10,11)$.
  • Figure 4: Training a $3 \times 15$ 2D network to classify a subset of MNIST digits. Each of the five $3 \times 3$ cells receives a PCA-processed input image. The output region $\mathcal{Y}$ contains five central nodes corresponding to digits ${0,1,3,6,9}$. Panel (a): data flow including PCA transformation, replication, and optical injection. Cells are separated by a blocking potential (red). (b, c): training and validation losses and accuracies; (d): confusion matrix with final accuracies on test set.
  • Figure 5: XOR task in (a) a 9-node 1D network with training restricted to $V$, fixed near-optimal pump weights $w$, and (d) a 7-node network with asymmetric input placement (sites 1, 3) and output at site 5. Top panels: training losses (b) for (a) and (e) for (d). Bottom: steady-state field magnitudes $|\Psi|^2$ for each input (c) and (f). (a)–(c) show that $V$--only training solves XOR for near-optimal fixed $w$; (d)–(f) show successful training despite broken spatial symmetry.
  • ...and 2 more figures