What Should We Ask Next? Retrieval-Aware Question Learning for Interactive ReID
cs.AI
Submitted: 2026-09-18
Updated: 2026-09-26
License: http://creativecommons.org/licenses/by/4.0/
The gist: Interactive retrieval under partial evidence is a sequential information-acquisition problem: an agent must decide which question will create the most useful evidence for the next retrieval update.
Terminology
Abstract
Interactive retrieval under partial evidence is a sequential information-acquisition problem: an agent must decide which question will create the most useful evidence for the next retrieval update. Existing systems train this decision by imitating an offline ordering of candidate QA pairs, although question value is determined by the response it elicits and its downstream effect on retrieval. We establish that candidate discriminativeness and perceived usefulness provide weak supervision for this objective, then introduce RAVEL, a retrieval-aware online reinforcement learning framework for interactive person re-identification. RAVEL initializes from supervised question generation, observes the current Top-4 candidates directly, and optimizes the question policy with rank feedback from the full question-answer-retrieval loop. Experiments on Interactive-PEDES show that RAVEL delivers progressively stronger retrieval performance across five interaction rounds. Further analysis shows that RAVEL reallocates the questioning budget toward localized open-ended attributes, which provide more useful retrieval evidence and yield the largest gains on initially difficult queries.
Sources
- VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering
- Semantically Self-Aligned Network for Text-to-Image Part-aware Person Re-identification
- Dialog-based Interactive Image Retrieval
- LLaVA-OneVision: Easy Visual Task Transfer
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
- Uncertainty-Aware Pedestrian Attribute Recognition via Evidential Deep Learning
- Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models
- Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Harnessing Weak Pair Uncertainty for Text-based Person Search
- ProRAG: Process-Supervised Reinforcement Learning for Retrieval-Augmented Generation
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
- Qwen2.5 Technical Report
- Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
- Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification
- ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection