Ambient @ EgoProactive 2026: Proactive Egocentric Assistance with Visually Grounded Supervision

arXiv:2609.07099 · cs.CV, cs.AI · Submitted 2026-09-07 · Read on arXiv

cs.CV, cs.AI

Submitted: 2026-09-07

Updated: 2026-09-10

Comments: Winning solution technical report for the EgoProactive track of the ECCV 2026 Wearable AI Grand Challenge

License: http://creativecommons.org/licenses/by/4.0/

The gist: We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first in the large-model division and second in the <=2B division.

Terminology

Abstract

We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first in the large-model division and second in the <=2B division. The task requires a wearable assistant to decide after each eight-second segment of egocentric video whether to intervene or remain silent. Our approach has two main components. First, we reformulate intervention timing as single-token classification. Rather than generating either interrupt <utterance> or silent, the model predicts yes or no, and we derive the decision from the renormalised probabilities of these two tokens. This formulation improved macro-F1 by 0.249 and G-mean by 0.30 over free-form generation. Second, because labelled data were limited to the released validation set, we generated additional supervision using a tool-calling video agent that inspects each clip and assigns intervention timestamps. A narration-only alternative was four times larger and ten times cheaper, but transferred worse than supervision from an unrelated real corpus, suggesting that visual grounding is more important than annotation volume for this task.

Related papers