Cost-efficient Active Learning for Referring Image Segmentation and Grounding
cs.CV, cs.AI
Submitted: 2026-08-31
Updated: 2026-09-01
Comments: Accepted to EMNLP 2026 Findings
Code: https://github.com/junbum766/ALRIS
License: http://creativecommons.org/licenses/by/4.0/
The gist: Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that
Terminology
Abstract
Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regions from visually similar ones. We tackle this by formulating active learning (AL) for VG under the realistic setting where only raw images are available without accompanying text. Since ground-truth text is unavailable, sample selection must estimate which images contain ambiguous regions that would require discriminative referring expressions. To address this, we generate auxiliary region-text pairs using foundation models, and introduce Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates. It allows our method to prioritize images with strong cross-region competition, which are more informative due to their visual ambiguity. We also design a referring-expression annotation interface that helps annotators quickly focus on writing discriminative language with a few clicks. Experiments on RIS and REC benchmarks show that our AL framework consistently outperforms several AL baselines, while a user study shows up to 1.6X faster description labeling of ours.
Sources
- MetaBox+: A new Region Based Active Learning Method for Semantic Segmentation using Priority Maps
- ACTRESS: Active Retraining for Semi-supervised Visual Grounding
- On uncertainty estimation in active learning for image segmentation
- Active Learning for Visual Question Answering: An Empirical Study
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
- RESMatch: Referring Expression Segmentation in a Semi-Supervised Manner
- MosMedData: Chest CT Scans With COVID-19 Related Findings Dataset
- Diverse mini-batch Active Learning
- A Simple Baseline with Single-encoder for Referring Image Segmentation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models