CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
cs.AI, cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Project page: https://sundongwei.github.io/CoEvolve_Project
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning
- GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision
- ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- EGM: Efficient Visual Grounding Language Models
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding
- IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection