EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports
cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: Accepted by EMNLP 2026 main conference
Code: https://github.com/NYCU-NLP-Lab/EgoArgus
License: http://creativecommons.org/licenses/by/4.0/
The gist: VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help.
Terminology
Abstract
VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving open whether models can arbitrate between visual evidence and user-provided language when the two are helpful, irrelevant, or conflicting. We introduce EgoArgus, a human-annotated dataset for evaluating egocentric assistants on understanding and decision tasks in five dialogue-video daily scenarios. Our results demonstrate that it is still challenging for current VLMs as reliable egocentric assistants, which requires identifying which modality is trustworthy and deciding when intervention is warranted. Deeper analysis also shows that existing modality bias mitigation methods are quite restricted to enhance performance, providing insights to aid practioners into the deployment of current VLMs as daily assistants.
Sources
- VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios
- Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering