GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
cs.CV, cs.AI, cs.LG, cs.RO
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/groundingpi/GroundAnything
Project page: https://groundingpi.github.io/groundanything
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Pix2seq: A Language Modeling Framework for Object Detection
- M$^{6}$Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- PaddleOCR 3.0 Technical Report
- RynnBrain: Open Embodied Foundation Models
- Emerging Properties in Unified Multimodal Pretraining
- Seed1.5-VL Technical Report
- Vision as Unified Multimodal Generation
- Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Kimi K3: Open Frontier Intelligence
- Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Qwen3 Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models