Transferable Equivariant Quantum Circuits for TSP: Generalization Bounds and Empirical Validation
Monit Sharma, Hoong Chuin Lau
TL;DR
The paper tackles generalization in quantum reinforcement learning for the Traveling Salesman Problem by leveraging permutation-equivariant quantum circuits (EQCs) to enable zero-shot transfer from small to larger problem instances. It derives a formal transfer bound that decomposes risk into a source-generalization term and a task-dissimilarity penalty $\mathcal{D}_{n \rightarrow m}$, which itself splits into parametric and structural components, both scaling with the size gap $m-n$. The theory is complemented by empirical validation on Euclidean TSP benchmarks, where policies trained on small graphs exhibit strong zero-shot transfer and can be further improved through targeted fine-tuning, with permutation-equivariant models consistently outperforming non-equivariant baselines. The work demonstrates that embedding symmetry into quantum models yields scalable, transferable QRL solutions for symmetric combinatorial tasks and outlines practical guidelines for curriculum-style training and symmetry-preserving architectures.
Abstract
In this work, we address the challenge of generalization in quantum reinforcement learning (QRL) for combinatorial optimization, focusing on the Traveling Salesman Problem (TSP). Training quantum policies on large TSP instances is often infeasible, so existing QRL approaches are limited to small-scale problems. To mitigate this, we employed Equivariant Quantum Circuits (EQCs) that respect the permutation symmetry of the TSP graph. This symmetry-aware ansatz enabled zero-shot transfer of trained parameters from $n-$city training instances to larger m-city problems. Building on recent theory showing that equivariant architectures avoid barren plateaus and generalize well, we derived novel generalization bounds for the transfer setting. Our analysis introduces a term quantifying the structural dissimilarity between $n-$ and $m-$node TSPs, yielding an upper bound on performance loss under transfer. Empirically, we trained EQC-based policies on small $n-$city TSPs and evaluated them on larger instances, finding that they retained strong performance zero-shot and further improved with fine-tuning, consistent with classical observations of positive transfer between scales. These results demonstrate that embedding permutation symmetry into quantum models yields scalable QRL solutions for combinatorial tasks, highlighting the crucial role of equivariance in transferable quantum learning.
