Table of Contents
Fetching ...

General transformation neural networks: A class of parametrized functions for high-dimensional function approximation

Xiaoyang Wang, Yiqi Gu

TL;DR

This work proposes a novel class of neural network-like parametrized functions, i.e., general transformation neural networks (GTNNs), for high-dimensional approximation, and performs an approximation error analysis of GTNNs, presenting their universal approximation properties for continuous functions, error bounds for Barron-type functions and error bounds of deep architectures.

Abstract

We propose a novel class of neural network-like parametrized functions, i.e., general transformation neural networks (GTNNs), for high-dimensional approximation. Conventional deep neural networks sometimes perform less accurately on learning problems trained with gradient descent, especially when the target function is oscillatory. To improve accuracy, we generalize the neuron's affine transformation to a broader class of functions that can capture complex shapes and offer greater capacity. Specifically, we discuss three types of GTNNs in detail: the cubic, quadratic and trigonometric transformation neural networks (CTNNs, QTNNs and TTNNs). We perform an approximation error analysis of GTNNs, presenting their universal approximation properties for continuous functions, error bounds for Barron-type functions and error bounds of deep architectures. Several numerical examples of regression problems are presented, demonstrating that CTNNs/QTNNs/TTNNs achieve higher accuracy than conventional fully connected neural networks.

General transformation neural networks: A class of parametrized functions for high-dimensional function approximation

TL;DR

This work proposes a novel class of neural network-like parametrized functions, i.e., general transformation neural networks (GTNNs), for high-dimensional approximation, and performs an approximation error analysis of GTNNs, presenting their universal approximation properties for continuous functions, error bounds for Barron-type functions and error bounds of deep architectures.

Abstract

We propose a novel class of neural network-like parametrized functions, i.e., general transformation neural networks (GTNNs), for high-dimensional approximation. Conventional deep neural networks sometimes perform less accurately on learning problems trained with gradient descent, especially when the target function is oscillatory. To improve accuracy, we generalize the neuron's affine transformation to a broader class of functions that can capture complex shapes and offer greater capacity. Specifically, we discuss three types of GTNNs in detail: the cubic, quadratic and trigonometric transformation neural networks (CTNNs, QTNNs and TTNNs). We perform an approximation error analysis of GTNNs, presenting their universal approximation properties for continuous functions, error bounds for Barron-type functions and error bounds of deep architectures. Several numerical examples of regression problems are presented, demonstrating that CTNNs/QTNNs/TTNNs achieve higher accuracy than conventional fully connected neural networks.
Paper Structure (35 sections, 8 theorems, 123 equations, 11 figures, 7 tables)

This paper contains 35 sections, 8 theorems, 123 equations, 11 figures, 7 tables.

Key Result

Theorem 3.1

Suppose $\sigma\in C(\mathbb{R})$ is not a polynomial and $\psi(\cdot;\hat{\theta}_0)$ is a non-zero constant for some $\hat{\theta}_0\in\Theta$. Let $d\geq1$ and $\Omega \subset \mathbb{R}^d$ be a compact set. For any $f \in C(\Omega)$ and $\epsilon>0$, there exists $M\in\mathbb{Z}^+$ and $a,b \in\ for all $x \in \Omega$.

Figures (11)

  • Figure 2.1: The architecture of a three-layer GTNN with $L=3$ and $M_1=M_2=2$ (the black dashed lines mean the scalar multiplication, and the blue solid lines mean general operations implemented by $\psi$).
  • Figure 2.2: The architecture of a three-layer QTNN with $L=3$ and $M_1=M_2=2$ (the black dashed lines mean the scalar multiplication).
  • Figure 2.3: Graphs of the piecewise quadratic/cubic polynomial (with fixed equispaced grids) and the two-layer QTNN/CTNN (with adaptive trainable grids).
  • Figure 4.1: The target function, the trained FNN and QTNN in Case 1.1.
  • Figure 4.2: Graphs of the target function and the trained FNN and CTNN at $(x_1,0,0,0,0)$ in Case 1.2.
  • ...and 6 more figures

Theorems & Definitions (12)

  • Remark 2.1
  • Theorem 3.1
  • proof
  • Theorem 3.2
  • proof
  • Lemma 3.3
  • Lemma 3.4
  • Theorem 3.5
  • Corollary 3.6
  • Corollary 3.7
  • ...and 2 more