WeightBridge: An Efficient Weight Transfer Library for Reinforcement Learning
cs.DC, cs.LG, cs.NI
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 25 pages, 16 figures
Code: https://github.com/PrimeIntellect-ai/prime-rl
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale
- Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
- RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
- Better & Faster Large Language Models via Multi-token Prediction
- AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
- LoRA: Low-Rank Adaptation of Large Language Models
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
- StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training
- Muon is Scalable for LLM Training
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
- Kimi K2: Open Agentic Intelligence
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Kimi K3: Open Frontier Intelligence
- Qwen3 Technical Report
- Learn Hard Problems During RL with Reference Guided Fine-tuning
- RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing