Table of Contents
Fetching ...

Near-optimal Prediction Error Estimation for Quantum Machine Learning Models

Qiuhao Chen, Yuling Jiao, Yinan Li, Xiliang Lu, Jerry Zhijian Yang

TL;DR

This work analyzes the performance of quantum machine learning models when only finite training data is available, addressing the gap left by generalization-error bounds. It introduces a theoretical framework based on covering and packing numbers and Gaussian denoise problems to bound the prediction error of both data-reuploading PQCs and linear QML models, proving a near-optimal rate of $\tilde{O}(T/N)$ and a matching $\Omega(T/N)$ lower bound. The results extend to 2-local constructions and to data-reuploading schemes, yielding $\tilde{O}(n q T / N)$ prediction error, with numerical experiments on univariate function approximation and quantum phase recognition validating the theory. These findings offer practical guarantees for near-term quantum devices and inform the design of data-encoding circuits to respect data and hardware constraints, complementing existing generalization analyses with a tighter, task-relevant perspective on learning with quantum models.

Abstract

Understanding the theoretical capabilities and limitations of quantum machine learning (QML) models to solve machine learning tasks is crucial to advancing both quantum software and hardware developments. Similarly to the classical setting, the performance of QML models can be significantly affected by the limited access to the underlying data set. Previous studies have focused on proving generalization error bounds for any QML models trained on a limited finite training set. We focus on the optimal QML models obtained by training them on a finite training set and establish a tight prediction error bound in terms of the number of trainable gates and the size of training sets. To achieve this, we derive covering number upper bounds and packing number lower bounds for the data re-uploading QML models and linear QML models, respectively, which may be of independent interest. We support our theoretical findings by numerically simulating the QML strategies for function approximation and quantum phase recognition.

Near-optimal Prediction Error Estimation for Quantum Machine Learning Models

TL;DR

This work analyzes the performance of quantum machine learning models when only finite training data is available, addressing the gap left by generalization-error bounds. It introduces a theoretical framework based on covering and packing numbers and Gaussian denoise problems to bound the prediction error of both data-reuploading PQCs and linear QML models, proving a near-optimal rate of and a matching lower bound. The results extend to 2-local constructions and to data-reuploading schemes, yielding prediction error, with numerical experiments on univariate function approximation and quantum phase recognition validating the theory. These findings offer practical guarantees for near-term quantum devices and inform the design of data-encoding circuits to respect data and hardware constraints, complementing existing generalization analyses with a tighter, task-relevant perspective on learning with quantum models.

Abstract

Understanding the theoretical capabilities and limitations of quantum machine learning (QML) models to solve machine learning tasks is crucial to advancing both quantum software and hardware developments. Similarly to the classical setting, the performance of QML models can be significantly affected by the limited access to the underlying data set. Previous studies have focused on proving generalization error bounds for any QML models trained on a limited finite training set. We focus on the optimal QML models obtained by training them on a finite training set and establish a tight prediction error bound in terms of the number of trainable gates and the size of training sets. To achieve this, we derive covering number upper bounds and packing number lower bounds for the data re-uploading QML models and linear QML models, respectively, which may be of independent interest. We support our theoretical findings by numerically simulating the QML strategies for function approximation and quantum phase recognition.
Paper Structure (26 sections, 21 theorems, 99 equations, 8 figures)

This paper contains 26 sections, 21 theorems, 99 equations, 8 figures.

Key Result

Lemma 1

The expected prediction error (defined in (eq: prediction error)) of the optimal QML models on a given training set (defined as the minimizer of (eqn:empirical_risk))is at most $O(\sqrt{T\log T/{N})}$.

Figures (8)

  • Figure 1: Performance analysis of QML models. The QML process begins with an initial model (white bullet) and progresses to a learned QML model (green bullet) through a hybrid optimization procedure, such as variational quantum algorithms. The distance between the learned QML model and the optimal QML model on the training set (blue bullet) is referred to as the optimization error, which quantifies the performance of the optimization procedure. The prediction error quantifies the gap between the optimal QML model on the training set and the optimal QML model (red bullet), measuring the influence resulting from a limited training set. The approximation error, representing the distance between the optimal QML model and the target function (star bullet), reflects the ability of the hypothesis space generated by the QML model to approximate the target functions.
  • Figure 2: Curve fitting results about the prediction error. Curve fitting is performed using a single-qubit parameterized quantum circuit. The QML model takes the $x$-label of the sample data as input, yielding an output within $[-1, 1]$. (a) The number of trainable parameters is fixed to be $60$. The fitting results under different training data sizes are demonstrated. (b) The training data size is fixed to be $32$. The fitting results under different numbers of trainable parameters are plotted. (c) The average loss of the optimal QML model on a training data set scales linearly in $1/N$ and $T$. In this task, the average loss of the optimal QML model is considered to be 0 when $T\geq 45$. Consequently, the average loss of the optimal QML model on a training dataset serves as a reliable indicator of its prediction error.
  • Figure 3: Classification results about the prediction error. (a) The phase diagram of the Hamiltonian. The phase boundary points (blue and red stars) are extracted from numerical simulations guided by congQuantumConvolutionalNeural2019. The ground state indicates the paramagnetic phase if $(h_1, h_2)$ lies above the blue dotted curve. Conversely, the ground state exhibits the antiferromagnetic phase if it falls below the green dotted curve. The ground state indicates the SPT phase if it falls between the green and blue dotted curves. The background shading illustrates the output from the QCNN with $N=40$. (b) The average loss and prediction error of the optimal QML model on a training data set scales linearly in $1/N$.
  • Figure 4: A parameterized quantum circuit to represent $U_4$. Note that $\{\text{U}_4(\boldsymbol{\gamma}):\boldsymbol{\gamma}\in(0,2\pi]^{15}\}\cong U_4$.
  • Figure 5: A parameterized quantum circuit for two-qubit amplitude encoding. For any unit vector $\boldsymbol{\gamma}\in\mathbb{C}^4$, let $\ket{\boldsymbol{\gamma}}=\gamma_{0}\ket{00}+\gamma_1\ket{01}+\gamma_2\ket{10}+\gamma_3\ket{11}$. It holds that $\ket{\boldsymbol{\gamma}}=\text{V}_4(\boldsymbol{\gamma})^{\dagger}\ket{00}$. We specify the choice of $W_1,W_2,W_3\in U_2$: Let $\texttt{Unitary(a, b)}={1}/{\sqrt{|\texttt{a}|^2+|\texttt{b}|^2}}\left(\texttt{a}\texttt{b}-\texttt{b}^*\texttt{a}^*\right)$ be a $2\times 2$ unitary matrix encoded with parameters $\texttt{a}$ and $\texttt{b}$. For any unit vector $\boldsymbol{\gamma}\in\mathbb{C}^4$, set $A_1=[\gamma_0\space\gamma_1]^T$ and $A_2=[\gamma_2\space\gamma_3]^T$. If $A_1^{\dagger}A_2=0$, let $k={\|A_2\|_2}/{\|A_1\|_2}$; otherwise, let $k=-{A_1^{\dagger}A_2}/{|A_1^{\dagger}A_2|}\cdot{\|A_2\|_2}/{\|A_1\|_2}$. Set $W_1=\texttt{Unitary}(\gamma_3-k\gamma_1, \gamma_2^*-k\gamma_0^*)^T$. Let $\text{CZ}(\mathbb{I}\otimes W_1)\ket{\boldsymbol{\gamma}}=\epsilon_0\ket{00}+\epsilon_1\ket{01}+\epsilon_2\ket{10}+\epsilon_3\ket{11}$ then $W_2=\texttt{Unitary}(\epsilon_1^*, \epsilon_3^*)$. Let $(W_2\otimes \mathbb{I})\text{CZ}(\mathbb{I}\otimes W_1)\ket{\boldsymbol{\gamma}}=\varrho_0\ket{00}+\varrho_1\ket{01}+\varrho_2\ket{10}+\varrho_3\ket{11}$ then $W_3=\texttt{Unitary}(\varrho_0^*, \varrho_1^*)^T$.
  • ...and 3 more figures

Theorems & Definitions (33)

  • Lemma 1
  • proof
  • Theorem 1
  • Theorem 2
  • Proposition 1
  • Lemma 2: Markov's inequality vershynin2018high
  • Lemma 3: Bernstein's inequality for bounded random variables vershynin2018high
  • Corollary 2
  • proof
  • Lemma 4: Hoeffding's inequality for general bounded random variables vershynin2018high
  • ...and 23 more