Libra: Efficient Resource Management for Agentic RL Post-Training
cs.LG, cs.AI, cs.DC
Submitted: 2026-06-02
Updated: 2026-09-16
Comments: 20 pages, 12 figures
Code: https://github.com/NVIDIA/MegatronLM
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
- Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- Echo: Simulating Distributed Training At Scale
- Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
- RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
- ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
- RollArt: Disaggregated Multi-Task Agentic RL Training at Scale
- AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
- History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- LLM Post-Training: A Deep Dive into Reasoning Large Language Models
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Training language models to follow instructions with human feedback
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks