Table of Contents
Fetching ...

Surface code scaling on heavy-hex superconducting quantum processors

Arian Vezvaee, Cesar Benito, Mario Morford-Oberst, Alejandro Bermudez, Daniel A. Lidar

TL;DR

This work tackles subthreshold surface-code scaling on IBM heavy-hex processors by co-designing a depth-minimizing SWAP-based embedding with bridge ancillas and robust dynamical decoupling, enabling anisotropic distance growth with $(d_x,d_z)=(3,5)$ and $(5,3)$. It introduces an entanglement-fidelity metric as a fit-free, SPAM-aware benchmark to assess subthreshold scaling, revealing that, under current hardware, there is no state-independent subthreshold scaling, though basis-specific gains arise and Willlow data suggests potential subthreshold scaling with hardware improvements. The study emphasizes that DD is essential to suppress coherent and non-Markovian noise and that suppression-factor metrics can mislead unless paired with careful DD optimization and EF-based analysis. The results provide a practical pathway for robust tests of subthreshold scaling on non-native architectures and quantify hardware targets (e.g., ~30% noise reduction and larger devices hosting $(5,5)$ codes) needed to realize genuine subthreshold operation. Overall, the work establishes EF as a principled benchmark for subthreshold scaling and guides hardware and control-layer improvements toward fault-tolerant quantum computation.

Abstract

Demonstrating subthreshold scaling of a surface-code quantum memory on hardware whose native connectivity does not match the code remains a central challenge. We address this on IBM heavy-hex superconducting processors by co-designing the code embedding and control: a depth-minimizing SWAP-based "fold-unfold" embedding that uses bridge ancillas, together with robust, gap-aware dynamical decoupling (DD). On Heron-generation devices we perform anisotropic scaling from a uniform distance 3 code to anisotropic distance (dx,dz) = (3,5) and (5,3) codes. We find that increasing dz (dx) improves the protection of Z-basis (X-basis) logical states across multiple quantum error correction cycles. Even if global subthreshold code scaling for arbitrary logical initial states is not yet achieved, we argue that it is within reach with minor hardware improvements. We show that DD plays a major role: it suppresses coherent ZZ crosstalk and non-Markovian dephasing that accumulate during idle gaps on heavy-hex layouts, and it eliminates spurious subthreshold claims that arise when scaled codes without DD are compared against smaller codes with DD. To quantify performance, we derive an entanglement fidelity metric that is computed directly from X- and Z-basis logical-error data and provides per-cycle, SPAM-aware bounds. The entanglement fidelity metric reveals that widely used single-parameter fits used to compute suppression factors can mischaracterize or obscure code performance when their assumptions are violated; we identify the strong assumptions of stationarity, unitality, and negligible logical SPAM required for those fits to be valid and show that they do not hold for our data. Our results establish a concrete path to robust tests of subthreshold surface-code scaling under biased, non-Markovian noise by integrating QEC with optimized DD on non-native architectures.

Surface code scaling on heavy-hex superconducting quantum processors

TL;DR

This work tackles subthreshold surface-code scaling on IBM heavy-hex processors by co-designing a depth-minimizing SWAP-based embedding with bridge ancillas and robust dynamical decoupling, enabling anisotropic distance growth with and . It introduces an entanglement-fidelity metric as a fit-free, SPAM-aware benchmark to assess subthreshold scaling, revealing that, under current hardware, there is no state-independent subthreshold scaling, though basis-specific gains arise and Willlow data suggests potential subthreshold scaling with hardware improvements. The study emphasizes that DD is essential to suppress coherent and non-Markovian noise and that suppression-factor metrics can mislead unless paired with careful DD optimization and EF-based analysis. The results provide a practical pathway for robust tests of subthreshold scaling on non-native architectures and quantify hardware targets (e.g., ~30% noise reduction and larger devices hosting codes) needed to realize genuine subthreshold operation. Overall, the work establishes EF as a principled benchmark for subthreshold scaling and guides hardware and control-layer improvements toward fault-tolerant quantum computation.

Abstract

Demonstrating subthreshold scaling of a surface-code quantum memory on hardware whose native connectivity does not match the code remains a central challenge. We address this on IBM heavy-hex superconducting processors by co-designing the code embedding and control: a depth-minimizing SWAP-based "fold-unfold" embedding that uses bridge ancillas, together with robust, gap-aware dynamical decoupling (DD). On Heron-generation devices we perform anisotropic scaling from a uniform distance 3 code to anisotropic distance (dx,dz) = (3,5) and (5,3) codes. We find that increasing dz (dx) improves the protection of Z-basis (X-basis) logical states across multiple quantum error correction cycles. Even if global subthreshold code scaling for arbitrary logical initial states is not yet achieved, we argue that it is within reach with minor hardware improvements. We show that DD plays a major role: it suppresses coherent ZZ crosstalk and non-Markovian dephasing that accumulate during idle gaps on heavy-hex layouts, and it eliminates spurious subthreshold claims that arise when scaled codes without DD are compared against smaller codes with DD. To quantify performance, we derive an entanglement fidelity metric that is computed directly from X- and Z-basis logical-error data and provides per-cycle, SPAM-aware bounds. The entanglement fidelity metric reveals that widely used single-parameter fits used to compute suppression factors can mischaracterize or obscure code performance when their assumptions are violated; we identify the strong assumptions of stationarity, unitality, and negligible logical SPAM required for those fits to be valid and show that they do not hold for our data. Our results establish a concrete path to robust tests of subthreshold surface-code scaling under biased, non-Markovian noise by integrating QEC with optimized DD on non-native architectures.
Paper Structure (37 sections, 2 theorems, 121 equations, 26 figures, 4 tables, 1 algorithm)

This paper contains 37 sections, 2 theorems, 121 equations, 26 figures, 4 tables, 1 algorithm.

Key Result

Lemma 1

For a given PTM, the entanglement fidelity depends only on the $3\times3$ block

Figures (26)

  • Figure 1: (a) Schematic of the surface codes implemented in this work on IBM's Heron-generation QPUs. The red boundary defines the $(3,5)$ surface code. The yellow boundary shows one of the possible sublattice $(3,3)$ codes that fits within the larger code. The other two $(3,3)$ sublattices are not shown. The green boundary defines the $(5,3)$ surface code. Similarly, there are three $(3,3)$ sublattices that fit within the $(5,3)$ code (not shown). (b) Stabilizer measurement construction for a heavy-hex surface code. The circuit combines the middle-out and traditional ancilla-based syndrome extraction schemes McEwen2023quantum, adapted to the heavy-hex connectivity Benito2025quantum. Next-nearest-neighbor CNOTs are implemented via bridge qubits using SWAP gates, and the resultant circuit is then simplified. Bridge qubits are not measured. Ancilla qubits are never reset; instead, we apply software processing to measurement outcomes depending on previous measurements. The bridge ancilla is not measured; all measurements are on the check ancilla. (c) $N=0$: logical state preparation via two stabilizer rounds, starting from the physical $\ket{0}^n$ state, followed by a measurement of data qubits. The goal of this cycle is SPAM calibration. (d) $N\ge 1$: QEC cycles, each comprising two stabilizer rounds, followed by a measurement of all the data qubits at the end of the $N$'th cycle. Only half the stabilizers are measured in parallel; a full cycle is two rounds.
  • Figure 2: Top: Logical error probability $\overline{p}^{\boldsymbol{d}}_{N,\alpha}$ versus QEC cycles $N$, for (a) $\overline{\ket{+}}$ and (b) $\overline{\ket{0}}$, including optimized DD. Open blue markers (“subs.”) show the three $(3,3)$ sublattices; filled blue markers (“avg.”) show their arithmetic mean. Note that the $(3,5)$ and $(5,3)$ codes contain different physical $(3,3)$ sublattices, hence the sublattice points differ between (a) and (b). Shaded regions and error bars denote $2\sigma$ confidence intervals. With respect to these intervals, both $(3,5)$ and $(5,3)$ outperform the average $(3,3)$ code, and for the first few cycles also outperform the best individual $(3,3)$ sublattice. Overall, $(3,5)$ shows the strongest improvement. Solid curves are fits to \ref{['eq:3param-model']}. Bottom: Cumulative differences $\delta^{\boldsymbol{d}}_x=\sum_{N}(\overline{p}^{\boldsymbol{d}}_{N,-}-\overline{p}^{\boldsymbol{d}}_{N,+})$ vs. $\delta^{\boldsymbol{d}}_z=\sum_{N}(\overline{p}^{\boldsymbol{d}}_{N,1}-\overline{p}^{\boldsymbol{d}}_{N,0})$ for (c) $(3,5)$ and (d) $(5,3)$, along with the average over their respective $(3,3)$ sublattices. Deviations from zero witness non‑unital logical noise, growing with $N$ and more pronounced for the $Z$ eigenstates.
  • Figure 3: Entanglement infidelity $1-F_{\mathrm{e}}(N)$ for (a) $(3,5)$ and (b) $(5,3)$. As in \ref{['fig:err-prob']}, blue open markers are the two sets of three $(3,3)$ sublattices relevant to each scaled code. Filled blue: arithmetic averages over those sublattices. We set $1-F_{\mathrm{e}}(0)=0$ by removing logical SPAM; $F_{\mathrm{e}}(N)$ is computed from $X/Z$ data using \ref{['eq:ent-fidelity']}.
  • Figure 4: Optimized EF metric $\Lambda_F$ [\ref{['eq:EF-metric']}] relative to the $(3,3)$ sublattice average (black with error bars) and relative to the best $(3,3)$ sublattice (purple, shaded). Also shown is the unoptimized EF metric $\frac{1-F_{\mathrm{e}}^{\boldsymbol{d}'}(N)}{1-F_{\mathrm{e}}^{\boldsymbol{d}}(N)}$ without DD (green, shaded). All results are for the $(5,3)$ code and its sublattices implemented on (a) ibm_aachen and (b) ibm_marrakesh. Values $>1$ would indicate basis‑independent subthreshold scaling. Consistent with \ref{['fig:ent-fidelity']}, this is not observed for ibm_aachen once DD is considered. For ibm_marrakesh we observe the onset of subthreshold scaling for $N\ge 5$ in terms of both the sublattice-average and the best sublattice, but with low statistical confidence. The no DD $(5,3)$ case exceeds $1$ for both ibm_aachen and ibm_marrakesh but is spurious, since the correct reference compares DD‑optimized $\max_{\{\text{DD,noDD}\}}F_{\mathrm{e}}$ for each code, as per \ref{['eq:EF-metric']}. To select the best sublattice, we compute $\frac{1}{9}\sum_{N=1}^{9}F_{\mathrm{e},s}^{(3,3)}(N)$ and maximize over $s\in\{1,2,3\}$.
  • Figure 5: The DD improvement factor in terms of (a) ratios of logical error probabilities and (b) entanglement infidelities. In more detail, in (a) we plot $\overline{p}^{\boldsymbol{d},\mathrm{noDD}}_{N,\alpha}/\overline{p}^{\boldsymbol{d}}_{N,\alpha}$, for the scaled codes and their respective $(3,3)$ code sublattice averages. A value above $1$ indicates that DD helps. The shaded confidence intervals are for the $\overline{\ket{+}}$ state and the error bars for the $\overline{\ket{0}}$ state. DD results in a significant improvement in the $\overline{\ket{+}}$ case, and has essentially no effect in the $\overline{\ket{0}}$ case. Results for $\overline{\ket{-}}$ (significant improvement) and $\overline{\ket{1}}$ (no improvement) are nearly identical and are not shown. In (b) we plot $\frac{1-F_{\mathrm{e}}^{\boldsymbol{d},\text{noDD}}(N)}{1-F_{\mathrm{e}}^{\boldsymbol{d}}(N)}$, which is the SPAM-free version of (a), and also simultaneously accounts for all four basis states. The $(3,5)$ code (green) is improved more by DD than the $(5,3)$ code (brown), while the opposite is true for their respective $(3,3)$ codes.
  • ...and 21 more figures

Theorems & Definitions (4)

  • Lemma 1
  • proof
  • Lemma 2
  • proof