AlphaQuanter: An End-to-End Tool-Orchestrated Agentic Reinforcement Learning Framework for Stock Trading
Zheye Deng, Jiashu Wang
TL;DR
AlphaQuanter tackles the core limitations of LLM-based trading by introducing a single-agent, tool-augmented RL framework that end-to-end optimizes a transparent decision workflow. It frames stock trading as a tool-augmented MDP and trains the agent to actively acquire information and selectively use tools, yielding interpretable reasoning traces and robust, profitable policies. Empirical results show state-of-the-art performance in ARR and SR across multiple stocks, with the 7B variant delivering consistent improvements and rich information-seeking behavior. The work highlights the value of optimizing the entire decision process (not just final predictions) for robust, auditable, and human-trader-friendly automated trading systems.
Abstract
While Large Language Model (LLM) agents show promise in automated trading, they still face critical limitations. Prominent multi-agent frameworks often suffer from inefficiency, produce inconsistent signals, and lack the end-to-end optimization required to learn a coherent strategy from market feedback. To address this, we introduce AlphaQuanter, a single-agent framework that uses reinforcement learning (RL) to learn a dynamic policy over a transparent, tool-augmented decision workflow, which empowers a single agent to autonomously orchestrate tools and proactively acquire information on demand, establishing a transparent and auditable reasoning process. Extensive experiments demonstrate that AlphaQuanter achieves state-of-the-art performance on key financial metrics. Moreover, its interpretable reasoning reveals sophisticated strategies, offering novel and valuable insights for human traders. Our code for data acquisition and agent training is publicly available at: https://github.com/AlphaQuanter/AlphaQuanter
