Learning to Zoom Efficiently with a Contrastive Curriculum
cs.CV, cs.CL
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: EMNLP 2026
Code: https://github.com/UKPLab/emnlp2026-zoom-in
License: http://creativecommons.org/licenses/by/4.0/
The gist: Using a zoom-in tool is an important foundational part of modern visual agents, because it allows to efficiently handle tasks involving high-resolution images.
Terminology
Abstract
Using a zoom-in tool is an important foundational part of modern visual agents, because it allows to efficiently handle tasks involving high-resolution images. Most previous methods need an extensive warm-start supervised fine-tuning phase for teaching models zoom-in. We show that this is not necessary by proposing a new intrinsic reward for learning tool use in MLLMs without the need for additional labels or warm-start SFT. Our InfoNCE-style reward uses a curriculum of increasingly hard negative tool calls as a contrastive training signal. Empirical experiments on V*, HRBench and MME-RealWorld show that our approach is competitive while being more efficient. When used as a drop-in replacement for SFT, we even outperform all baselines. To directly measure the zoom-in ability of models, we further introduce the scalable synthetic Muffin&Chihuahua (M&C) dataset. Each image consists of a grid with every cell either showing a muffin or chihuahua. Leveraging the M&C dataset's unique region of interest labels, we find that recall is the metric that most strongly correlates the zoom-in region with final task performance. Our model and code for reproduction is publicly available under https://github.com/UKPLab/emnlp2026-zoom-in
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- In Defense of the Triplet Loss for Person Re-Identification
- Representation Learning with Contrastive Predictive Coding
- Needles in Haystacks: On Classifying Tiny Objects in Large Images
- Exploration in Deep Reinforcement Learning: A Survey
- Empowerment -- an Introduction
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- The Invisible Leash: Why RLVR May or May Not Escape Its Origin
- MMSearch-R1: Incentivizing LMMs to Search
- Gemma 4 Technical Report
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Fine-Tuning Language Models from Human Preferences
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models