PIVOT: Pivot-Aware On Policy Self Distillation for Multi-Turn VLM Agents
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/mpSchrader/gym-sokoban
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- AI2-THOR: An Interactive 3D Environment for Visual AI
- Self-Distilled Agentic Reinforcement Learning
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling
- Solving math word problems with process- and outcome-based feedback
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
- SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models