Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
cs.CL
Submitted: 2026-06-16
Updated: 2026-09-26
Code: https://github.com/Deep-Agent/R1-V
Terminology
Sources
- GPT-4o System Card
- Gemini: A Family of Highly Capable Multimodal Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Understanding R1-Zero-Like Training: A Critical Perspective
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
- The Art of Scaling Reinforcement Learning Compute for LLMs
- SmolVLM: Redefining small and efficient multimodal models
- Distilling the Knowledge in a Neural Network
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- DistiLLM: Towards Streamlined Distillation for Large Language Models
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- A Survey of On-Policy Distillation for Large Language Models
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Fast and Effective On-policy Distillation from Reasoning Prefixes
- Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering