Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/vermouth599/Spatial-OPSD
Terminology
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- Qwen3-VL Technical Report
- Scaling Spatial Intelligence with Multimodal Foundation Models
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
- SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
- 4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
- Emerging Properties in Unified Multimodal Pretraining
- Gemma 3 Technical Report
- MiniLLM: On-Policy Distillation of Large Language Models
- Reinforcement Learning via Self-Distillation
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
- Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
- Visual-Advantage On-Policy Distillation for Vision-Language Models
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- 3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
- SAM 2: Segment Anything in Images and Videos
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models