Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
cs.LG, cs.CL, cs.DC
Submitted: 2026-07-22
Updated: 2026-09-22
Comments: update tech report
DOI: 10.13140/RG.2.2.23375.65447
Code: https://github.com/NVIDIA-NeMo/labs-molt
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agentic reinforcement learning requires infrastructure that researchers can modify without sacrificing model scale or control over agent execution.
Terminology
Abstract
Agentic reinforcement learning requires infrastructure that researchers can modify without sacrificing model scale or control over agent execution. We present Molt, a lightweight PyTorch-native framework that combines trillion-parameter training with standard agent interfaces. Molt integrates four capabilities: a compact training implementation built on composable model parallelism; unified OpenAI and Anthropic interfaces with automatic trajectory segmentation after context compaction; fully asynchronous rollout and optimization; and distributed experience storage for long, multimodal trajectories. Existing agents retain their execution and context-management logic while a shared capture layer records generated tokens and behavior probabilities. Rollout workers place heavy experience payloads in Ray's object store, and trainer ranks retrieve their assigned experiences by reference, avoiding a centralized gather of the full rollout batch. The framework-owned RL implementation comprises approximately 9.2K Python code lines, and its rollout, weight-refit, and training-update path has executed end to end on a one-trillion-parameter policy. On a 35B multimodal mixture-of-experts workload, speculative decoding accelerates the generation stage by 5.14x, and optimizer offload reduces peak actor memory by 18.3 GB. Together, these results establish a compact training framework for agentic RL research at trillion-parameter scale.
Sources
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- OpenAI Gym
- RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning
- DeepSeek-V3 Technical Report
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- RollArt: Disaggregated Multi-Task Agentic RL Training at Scale
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
- DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
- Understanding R1-Zero-Like Training: A Critical Perspective
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Laminar: A Scalable Asynchronous RL Post-Training Framework
- Polar: Agentic RL on Any Harness at Scale
- Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks