Plasma Shape Control via Zero-shot Generative Reinforcement Learning
Niannian Wu, Rongpeng Li, Zongyu Yang, Yong Xiao, Ning Wei, Yihang Chen, Bo Li, Zhifeng Zhao, Wulyu Zhong
TL;DR
This work tackles the limited adaptability of PID control and generalization gaps in RL for tokamak plasma shape control by proposing a zero-shot framework that learns a foundation policy from a large offline PID dataset. The method combines Generative Adversarial Imitation Learning with a Hilbert space latent encoder to produce a goal-conditioned policy that can track diverse plasma configurations without task-specific retraining, trained in a data-driven HL-3 dynamics environment. Key contributions include the Hilbert encoder that encodes target states into a navigable latent space, an adversarial discriminator to enforce PID-like behavior, and a PPO-based actor-critic that optimizes a hybrid imitation/directed-reward objective, demonstrated by precise, stable zero-shot tracking across multiple current and shape scenarios in high-fidelity simulations. The results suggest a data-efficient path toward flexible intelligent plasma control suitable for future fusion devices, reducing the need for retraining when targets or operating conditions change.
Abstract
Traditional PID controllers have limited adaptability for plasma shape control, and task-specific reinforcement learning (RL) methods suffer from limited generalization and the need for repetitive retraining. To overcome these challenges, this paper proposes a novel framework for developing a versatile, zero-shot control policy from a large-scale offline dataset of historical PID-controlled discharges. Our approach synergistically combines Generative Adversarial Imitation Learning (GAIL) with Hilbert space representation learning to achieve dual objectives: mimicking the stable operational style of the PID data and constructing a geometrically structured latent space for efficient, goal-directed control. The resulting foundation policy can be deployed for diverse trajectory tracking tasks in a zero-shot manner without any task-specific fine-tuning. Evaluations on the HL-3 tokamak simulator demonstrate that the policy excels at precisely and stably tracking reference trajectories for key shape parameters across a range of plasma scenarios. This work presents a viable pathway toward developing highly flexible and data-efficient intelligent control systems for future fusion reactors.
