SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

Zhiyu Mei; Wei Fu; Jiaxuan Gao; Guangju Wang; Huanchen Zhang; Yi Wu

SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

Zhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang, Huanchen Zhang, Yi Wu

TL;DR

SRL tackles the challenge of scaling reinforcement learning to large-scale clusters by introducing a general dataflow abstraction built on workers, streams, and services. It decouples actor (environment), policy (inference), and trainer (training) workloads into three core worker types and adds dataflow mechanisms (inference and sample streams) plus a parameter server to enable massively parallel data generation and training across heterogeneous hardware. Empirically, SRL delivers superior end-to-end training throughput compared with open-source baselines and scales effectively to clusters with tens of thousands of CPU cores, while also matching or accelerating learning performance on common RL benchmarks and in the challenging Hide-and-Seek environment (up to 5x wall-clock speedups with GPU inference). The work demonstrates substantial practical impact by enabling large-scale, flexible RL experimentation and rapid algorithm development, surpassing prior academic systems in throughput and scalability and approaching production-scale capabilities.

Abstract

The ever-growing complexity of reinforcement learning (RL) tasks demands a distributed system to efficiently generate and process a massive amount of data. However, existing open-source libraries suffer from various limitations, which impede their practical use in challenging scenarios where large-scale training is necessary. In this paper, we present a novel abstraction on the dataflows of RL training, which unifies diverse RL training applications into a general framework. Following this abstraction, we develop a scalable, efficient, and extensible distributed RL system called ReaLlyScalableRL, which allows efficient and massively parallelized training and easy development of customized algorithms. Our evaluation shows that SRL outperforms existing academic libraries, reaching at most 21x higher training throughput in a distributed setting. On learning performance, beyond performing and scaling well on common RL benchmarks with different RL algorithms, SRL can reproduce the same solution in the challenging hide-and-seek environment as reported by OpenAI with up to 5x speedup in wall-clock time. Notably, SRL is the first in the academic community to perform RL experiments at a large scale with over 15k CPU cores. SRL source code is available at: https://github.com/openpsi-project/srl .

SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

TL;DR

Abstract

Paper Structure (41 sections, 12 figures, 12 tables)

This paper contains 41 sections, 12 figures, 12 tables.

Introduction
Background & Motivation
Reinforcement Learning System
Limitations of Existing Systems
System Design & Architecture
High-Level Design of SRL
System Components & Implementation
Performance Optimization
User-friendly and Extensible Designs
Experiments
Training Throughput
Environments & Algorithm
Comparison with Baselines
Large-Scale Architecture Evaluation
Learning Performance
...and 26 more sections

Figures (12)

Figure 1: Capabilities of open-source distributed RL systems.
Figure 1: IMPALA-style (left) and SEED-style (right) architecture implementations on a cluster with GPU nodes. The former merges environment simulation and policy inference in a single CPU/GPU node, while the latter merges policy inference and training on centralized GPU node. Note that in SEED-style, GPU nodes running environment simulation rely on the training GPU node for policy inference, while in IMPALA-style, they rely on local CPU/GPU for policy inference.
Figure 2: (left) In SRL abstraction, workers host task handlers to execute computing tasks. Workers are connected by data streams and supported by services. (middle) Based on the abstraction, the architecture for a typical RL workflow in SRL incorporates 3 types of core workers, 2 types of streams and the parameter services. (right) In an execution instance of SRL, workers are assigned appropriate resources on heterogeneous nodes in a distributed cluster. Data streams exploit fastest available communication substrates to ensure high-throughput data transmission.
Figure 2: Training throughput with 8 A100 GPU trainers with distributed actors. # CPU Cores (peak): CPU cores used for training sample generation when trainers reaches peak performance.
Figure 3: Training FPS of SRL and baselines on a single machine. SeedRL with 32 and 64 CPU cores results in GPU out-of-memory. SeedRL and Rlpyt do not support multi-agent environments.
...and 7 more figures

SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

TL;DR

Abstract

SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

Authors

TL;DR

Abstract

Table of Contents

Figures (12)