Table of Contents
Fetching ...

Parameter Analysis and Optimization of Layer Fidelity for Quantum Processor Benchmarking at Scale

Maria Jose Lozano Palacio, Hasan Nayfeh, Matthew Ware, David C. McKay

TL;DR

This paper advances scalable quantum processor benchmarking by refining the layer fidelity (LF) framework into a practical protocol for identifying optimal N-qubit chains. It combines isolated RB and grid-layer fidelity to select high-performing chains using a cost function $LF = \prod_{j} F_j$ with $F_j = \frac{(1 - \epsilon_j)(d+1) - 1}{d}$ and $d \in \{2,4\}$, then evaluates the best chains with EPLG = $\frac{4}{5}(1 - LF^{N-1})$, achieving 40–70% lower EPLG than random chains on IBM devices. EPLG is demonstrated as a robust stability metric over 100 days, capable of signaling both edge-localized and device-wide degradation, including TLS-induced transients. The study also analyzes RB-fit parameter sensitivity and shows that longer 2Q gate durations can markedly increase EPLG on Eagle R3 architectures, providing practical guidelines for large-scale benchmarking and potential generalization to other layer topologies.

Abstract

With the continued scaling of quantum processors, holistic benchmarks are essential for extensively evaluating device performance. Layer fidelity is a benchmark well-suited to assessing processor performance at scale. Key advantages of this benchmark include its natural alignment with randomized benchmarking (RB) procedures, crosstalk awareness, fast measurements over large numbers of qubits, high signal-to-noise ratio, and fine-grained information. In this work, we extend the analysis of the original layer fidelity manuscript to optimize parameters of the benchmark and extract deeper insights of its application. We present a robust protocol for identifying optimal qubit chains of length N, demonstrating that our method yields error per layered gate (EPLG) values 40%-70% lower than randomly selected chains. We further establish layer fidelity as an effective performance monitoring tool, capturing both edge-localized and device-wide degradation by tracking optimal chains of length 50 and 100, and fixed chains of length 100. Additionally, we refine error analysis by proposing parameter bounds on the number of randomizations and Clifford lengths used in direct RB fits, minimizing fit uncertainties. Finally, we analyze the impact of varying gate durations on layer fidelity measurements, showing that prolonged gate times leading to idling times significantly increase layered two-qubit (2Q) errors on Eagle R3 processors. Notably, we observe a 95% EPLG increase on a fixed chain in an Eagle R3 processor when some gate durations are extended by 65%. These findings extend the applicability of the layer fidelity benchmark and provide practical guidelines for optimizing quantum processor evaluations.

Parameter Analysis and Optimization of Layer Fidelity for Quantum Processor Benchmarking at Scale

TL;DR

This paper advances scalable quantum processor benchmarking by refining the layer fidelity (LF) framework into a practical protocol for identifying optimal N-qubit chains. It combines isolated RB and grid-layer fidelity to select high-performing chains using a cost function with and , then evaluates the best chains with EPLG = , achieving 40–70% lower EPLG than random chains on IBM devices. EPLG is demonstrated as a robust stability metric over 100 days, capable of signaling both edge-localized and device-wide degradation, including TLS-induced transients. The study also analyzes RB-fit parameter sensitivity and shows that longer 2Q gate durations can markedly increase EPLG on Eagle R3 architectures, providing practical guidelines for large-scale benchmarking and potential generalization to other layer topologies.

Abstract

With the continued scaling of quantum processors, holistic benchmarks are essential for extensively evaluating device performance. Layer fidelity is a benchmark well-suited to assessing processor performance at scale. Key advantages of this benchmark include its natural alignment with randomized benchmarking (RB) procedures, crosstalk awareness, fast measurements over large numbers of qubits, high signal-to-noise ratio, and fine-grained information. In this work, we extend the analysis of the original layer fidelity manuscript to optimize parameters of the benchmark and extract deeper insights of its application. We present a robust protocol for identifying optimal qubit chains of length N, demonstrating that our method yields error per layered gate (EPLG) values 40%-70% lower than randomly selected chains. We further establish layer fidelity as an effective performance monitoring tool, capturing both edge-localized and device-wide degradation by tracking optimal chains of length 50 and 100, and fixed chains of length 100. Additionally, we refine error analysis by proposing parameter bounds on the number of randomizations and Clifford lengths used in direct RB fits, minimizing fit uncertainties. Finally, we analyze the impact of varying gate durations on layer fidelity measurements, showing that prolonged gate times leading to idling times significantly increase layered two-qubit (2Q) errors on Eagle R3 processors. Notably, we observe a 95% EPLG increase on a fixed chain in an Eagle R3 processor when some gate durations are extended by 65%. These findings extend the applicability of the layer fidelity benchmark and provide practical guidelines for optimizing quantum processor evaluations.
Paper Structure (6 sections, 5 figures)

This paper contains 6 sections, 5 figures.

Figures (5)

  • Figure 1: (a) EPLG vs chain length for the optimal $N$-qubit chain (red) and 3 random chains (blue) on $ibm\_marrakesh$ (Heron R2). The lowest EPLG corresponds to the optimal chain, with $3.2 \times 10^{-3}$, followed by a random chain with $8 \times 10^{-3}$. (b) EPLG vs chain length for the optimal $N$-qubit chain (red) and 3 random chains (blue) on $ibm\_brisbane$ (Eagle R3). The lowest EPLG corresponds to the optimal chain, with $1.4 \times 10^{-2}$, followed by a random chain with $1.9 \times 10^{-2}$. (c) Horizontal grid configuration (red) and vertical grid configuration (blue) for a Heron R2 processor. (d) Frequency of occurrence of the optimal chain (isolated RB strategy (blue) vs grid strategy (red)) for a Heron R2 (left) and Eagle R3 (right). Red has a $60\%$ occurrence percentage for Heron R2 and a $70\%$ occurrence percentage for Eagle R3.
  • Figure 2: (a) EPLG vs time for $ibm\_fez$ (Heron R2, left) and $ibm\_torino$ (Heron R1, right) for chains of length 50 (blue) and 100 (red). Dashed lines indicate the high outlier region over the most recent 15 days. On day 62, the EPLG of the 100-qubit chain on $ibm\_torino$ increases to $1.2 \times 10^{-2}$. (b) EPLG vs chain length for $ibm\_torino$ on day 62 (blue) and day 65 (red). On day 62, there are two 2Q gates at lengths 78 and 79 with poor fidelities due to a TLS interaction involving the qubits in those gates. EPLG recovers by day 65.
  • Figure 3: (a) EPLG vs time for $ibm\_fez$ (Heron R2, blue) and $ibm\_torino$ (Heron R1, red) for a fixed chain of length N=100. On day 62, the EPLG of $ibm\_torino$ increases to $2.8 \times 10^{-2}$ due to a strong TLS interaction involving one qubit in the chain. (b) Normal quantile distribution of EPLG values over a 100-day period for fixed chains on $ibm\_fez$ (Heron R2, blue) and $ibm\_torino$ (Heron R1, red).
  • Figure 4: (a) EPLG vs number of randomizations on $ibm\_fez$ (Heron R2) for a chain of length 100. Shaded areas denote the standard error, propagated from error fits in RB decay. Nominal EPLG starts stabilizing at around $r=6$. Shaded EPLG areas start stabilizing at around $r=20$. (b) EPLG vs Clifford lengths on $ibm\_fez$ for a chain of length 100. Shaded areas denote the standard error, propagated from error fits in RB decay. Nominal EPLG starts stabilizing at around $100$ Cliffords. Shaded EPLG areas start stabilizing at around $300$ Cliffords.
  • Figure 5: (a) Layered RB (left) and isolated RB (right) circuits. Layered circuits contain an alignment barrier, unlike the isolated case. Isolated RB with a delay, not picture here, looks similar to the layered case, except that it has a separation of 2 idle qubits between 2Q gates. (b) Chain used to measure 2Q errors on $ibm\_marrakesh$. The chain for $ibm\_shebrooke$ is similar. (c) 2Q error distributions for $ibm\_marrakesh$. (Heron R2). Blue line is layered RB including long gates, light blue line is isolated RB with a delay, and red line is isolated RB. Distributions are fairly tight. (d) 2Q error distributions for $ibm\_sherbrooke$ (Eagle R3). The ratio between the median 2Q error of layered RB and isolated RB is 3.9x, while the ratio between isolated RB with and without a delay is 2x. (e) EPLG vs chain length for $ibm\_marrakesh$ for a fixed chain with 6 gates of varying gate lengths. Blue line corresponds to layer fidelity run on all gates at the maximal gate length of 80ns ($3.68 \times 10^{-3}$). Red line excludes all 6 gates with longer durations, and runs the rest at their corresponding duration of 68ns ($3.55 \times 10^{-3}$). (f) EPLG vs chain length for $ibm\_sherbrooke$ for a fixed chain with 7 gates of varying gate lengths. Blue line corresponds to layer fidelity run on all gates at the maximal gate length of 881ns ($3.7 \times 10^{-2}$). Red line excludes all 7 gates with longer durations, and runs the rest at their corresponding duration of 533ns ($1.9 \times 10^{-2}$).