ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation
cs.CV
Submitted: 2026-10-08
Updated: 2026-10-08
Terminology
Sources
- When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation
- Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception
- Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation
- TALENT: Target-aware Efficient Tuning for Referring Image Segmentation
- OpenVLA: An Open-Source Vision-Language-Action Model
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation
- Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- LaSagnA: Language-based Segmentation Assistant for Complex Queries
- Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
- MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
- RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization
- KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models