AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
cs.LG, cs.AI
Submitted: 2025-07-02
Updated: 2026-09-11
Code: https://github.com/huggingface/trl
Project page: https://nvidia.github.io/TensorRTLLM
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs).
Terminology
Abstract
Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs). Traditional task-collocated RL frameworks suffer from significant scalability bottlenecks, while task-separated RL frameworks face challenges in managing complex dataflows and resolving resource idling. Furthermore, most existing frameworks are tightly coupled with LLM training or inference engines, making them difficult to support custom-designed engines. To address these challenges, we propose AsyncFlow, an asynchronous streaming RL framework tailored for efficient post-training. Specifically, we introduce a distributed data storage and transfer module that provides panoramic data management and fine-grained scheduling capabilities in a fully streamed manner. This architecture inherently enables automated pipeline overlapping among RL tasks and dynamic load-balancing. Moreover, we propose an asynchronous producer-consumer workflow, which is engineered to minimize computational idleness by strategically deferring the parameter update process within staleness thresholds. Finally, the core capabilities of AsyncFlow are architecturally decoupled from underlying training and inference engines and encapsulated by service-oriented user interfaces, offering a modular and customizable user experience. Extensive experiments demonstrate an average throughput of 1.59x compared to the state-of-the-art baseline. The architecture presented in this work provides actionable insights for designing next-generation RL training systems.
Sources
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- Scaling Laws for Neural Language Models
- DeepSeek-V3 Technical Report
- Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
- Proximal Policy Optimization Algorithms
- Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
- NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
- HybridFlow: A Flexible and Efficient RLHF Framework
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Finetuned Language Models Are Zero-Shot Learners
- Qwen2.5 Technical Report
- DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- A Survey of Large Language Models
- StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
- Optimizing RLHF Training for Large Language Models with Stage Fusion
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks