HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training
cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Adacc: An Adaptive Framework Unifying Compression and Activation Recomputation for LLM Training
- Training Deep Nets with Sublinear Memory Cost
- OASIS: Online Activation Subspace Learning for Memory-Efficient Training
- DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
- The Llama 3 Herd of Models
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Schedule-Level Shared-Prefix Reuse for LLM RL Training
- Let's Verify Step by Step
- Cost-Aware Learning
- FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
- Qwen2.5 Technical Report
- Not all tokens are needed(NAT): token efficient reinforcement learning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models
- LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training
- Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
- Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models
- How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
- QUADS: Stabilizing NVFP4 Reinforcement Learning for MoE via QUantization-error Alignment across Dual Sides
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks