RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models
cs.CV, cs.AI, cs.IR
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Accepted by ECCV 2026. 22 pages, including references and appendix
Code: https://github.com/kakaobrain/coyo-dataset
License: http://creativecommons.org/licenses/by/4.0/
The gist: Region-level retrieval aims to align user-specified image regions with relevant regions or textual descriptions, playing a crucial role in realworld applications such as e-commerce product search and
Terminology
Abstract
Region-level retrieval aims to align user-specified image regions with relevant regions or textual descriptions, playing a crucial role in realworld applications such as e-commerce product search and RAG. Although recent Large Multimodal Models (LMMs) have made significant strides in multimodal retrieval, they primarily focus on global-level tasks and struggle to capture effective region-level representations. To bridge this gap, we present RegRet, an LMM-based Region-level Retrieval framework that enhances the regional representations without compromising overall global retrieval performance. At its core, RegRet integrates a Region-Aware Encoder to capture detailed regional features while balancing them with the global background context. To further enhance the fine-grained understanding and discriminability of representations, we design a multi-stage training pipeline that includes detailed localized captioning and regional contrastive learning tasks. In addition, considering the absence of region-level contrastive training data and the limited diversity of evaluation tasks in current benchmarks, we introduce the REGMB benchmark. It comprises 225k contrastive pairs, covering four multimodal retrieval tasks. Extensive experiments validate the effectiveness of our approach. RegRet outperforms strong baselines in the zero-shot setting. Further training with contrastive learning leads to an average improvement of more than 20% on both REGMB and public benchmarks, while achieving comparable or better results on global-level retrieval tasks.
Sources
- mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
- ColPali: Efficient Document Retrieval with Vision Language Models
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- E5-V: Universal Embeddings with Multimodal Large Language Models
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
- Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
- Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
- Segment Anything
- Region-based Contrastive Pretraining for Medical Image Retrieval with Anatomic Query
- Describe Anything: Detailed Localized Image and Video Captioning
- MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- Representation Learning with Contrastive Predictive Coding
- EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models