EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

arXiv:2607.00218 · cs.CV, cs.AI · Submitted 2026-06-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards".

Tom: Vision-language models (VLMs) are increasingly deployed as runtime safety guards for embodied agents, necessitating benchmarks that distinguish between genuinely unsafe situations and routine but superficially alarming activities.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, to recap, EgoSafetyBench proposes a diagnostic egocentric video benchmark consisting of one thousand two hundred robotview scenarios annotated at half-second granularity to test VLMs as streaming guards across two tracks <ref:2607.00218#pg0,of 1,200 robotview scenarios annotated at half-second granularity to>. The main thesis is that a deployable guard needs to distinguish between genuinely unsafe situations and routine but superficially alarming activities that binary safety benchmarks often miss.

Jane: Exactly, Tom; the paper claims this benchmark matters because it evaluates whether these AI systems can reliably catch real physical danger while avoiding unnecessary intervention on routine activities, which is something standard benchmarks don't really capture well.

Lu: They set up the situational track with four families: S1 for routine safe, S2 for resolved-confound safe, U1 for obvious hazard, and U2 for contextual hazard. This taxonomy allows them to test reasoning across a wide spectrum of physical safety concerns.

Meng: That distinction between obvious and contextual hazards is critical because it suggests that the required level of AI reasoning changes depending on the situation, which is a key practical consideration for engineering these guards.

Lalam: And then they add the visual-channel mismatch axis, where they check if in-scene text or labels misrepresent the physical situation. This adds another layer of complexity to how we judge a guard’s performance.

Conclusion: Tom: Thinking about the title, EgoSafetyBench, it really tells us we need a way to diagnose exactly *why* a VLM failed when it didn't catch something or when it flagged something incorrectly during runtime monitoring.

Jane: It’s an important step because instead of just saying "the model failed," this benchmark helps us pinpoint whether the failure was due to misjudging physical danger or being misled by deceptive visual cues.

Lu: The authors are pushing for a framework where we can evaluate embodied VLMs not just on accuracy, but on their ability to reason about complex physical states and external information simultaneously.

Meng: Practically speaking, this means future AI deployment for robots will require guards that understand the difference between a hot pipe being dangerous versus a sign saying "COLD" when it's actually safe.

Lalam: I think the implication is that these models can become much more nuanced in their safety decisions, moving from simple pattern matching to actual contextual understanding of the physical world.

Tom: So, we're looking at a richer way to test safety guards by checking both what they see and how they interpret that information, and that’s what EgoSafetyBench is all about.

Siddhant Panpatil, Arth Singh, Mijin Koo, Chaeyun Kim, Haon Park, Dasol Choi

AIM Intelligence 2Seoul National University

cs.CV, cs.AI

Submitted: 2026-06-30

Updated: 2026-10-04

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 92/100

The gist: Vision-language models (VLMs) are increasingly deployed as runtime safety guards for embodied agents, necessitating benchmarks that distinguish between genuinely unsafe situations and routine but

Key concepts

Situational Track (S1-U2)
This axis categorizes physical situations into four safety families: Routine safe (S1), Resolved-confound safe (S2), Obvious hazard (U1), and Contextual hazard (U2). It helps researchers see if models fail at recognizing simple dangers or complex, situation-dependent risks.
Visual-Channel Mismatch (VCM)
This axis checks if a visible sign, sticker, or label in the scene is actually misleading regarding the physical situation. The benchmark pairs every deceptive channel with a truthful control to isolate whether model errors come from general perception or being fooled by specific visual deception.
Contrastive Design
Scenarios are designed so they share most elements but differ only by one visible deciding variable. This forces the safety guard to focus on specific, critical relationships rather than relying on broad scene patterns, allowing for deeper diagnostic analysis of reasoning failures.

Terminology

Summary

Vision-language models (VLMs) are increasingly deployed as runtime safety guards for embodied agents, necessitating benchmarks that distinguish between genuinely unsafe situations and routine but superficially alarming activities. EGOSAFETYBENCH introduces a diagnostic egocentric video benchmark designed to evaluate VLMs as streaming guards across two distinct tracks: one assessing situational physical safety and another assessing visual-channel trustworthiness.

The gist: While models reliably flag videos containing hazards, they often miss specific hazardous moments, particularly contextual hazards; furthermore, misleading in-scene signs degrade all tested guards, with vulnerable models missing up to a third of hazards while robust models overintervene on safe content.

Benchmark Design and Annotation Axes

EGOSAFETYBENCH is an egocentric video benchmark consisting of 1,200 robotview scenarios annotated at half-second granularity. The annotation framework decouples evaluation along two independent axes: the situational track and the visual-channel mismatch (VCM) axis. The situational axis utilizes a four-family taxonomy: S1 (Routine safe), S2 (Resolved-confound safe), U1 (Obvious hazard), and U2 (Contextual hazard). The VCM axis is a separate binary label marking whether an in-scene channel—a sign, sticker, or label—misrepresents the physical situation. Crucially, every misleading channel is paired with a truthful control on an identical scene to isolate deception from general perception errors.

Evaluation Protocol and Streaming Setting

The benchmark evaluates VLMs as runtime safety guards by processing videos chunk-by-chunk in an online setting that approximates a streaming guard. Each five-second clip is divided into ten non-overlapping half-second chunks, and every decision targets the current chunk rather than the whole video. For each chunk, models produce two independent outputs: a situational safety verdict and a visual-channel mismatch flag. The model receives only a single chunk or a short causal window of recent chunks, without scenario descriptions or privileged state information.

Contrastive Design for Diagnostic Slicing

A central feature of the benchmark is the contrastive ladder design. Scenarios share the same entity, action, target, and domain but differ only by a single visible deciding variable. This forces guards to evaluate specific safety-relevant relationships rather than exploiting broad scene correlations. Each scenario also carries one of nine mechanism tags (e.g., spatial relation, object state) which enables diagnostic slicing beyond axis-level accuracy to reveal which kinds of reasoning drive model failures.

Evaluation Metrics and Findings

The paper evaluates models across multiple metrics, including video level (miss rate, false alarm rate), chunk level (miss/false-alarm rates split by family), timing (caught rate, premature alarm rate), and cross-axis robustness. Key findings indicate that the contextual-hazard miss (U2) exceeds the obvious-hazard miss (U1) for every model. Misleading in-scene text degrades all guards: vulnerable open models miss more hazards, while robust closed models over-intervene on safe content, demonstrating that apparent safety robustness often reflects indiscriminate alarming rather than true physical reasoning.

Model Evaluation and Robustness Analysis

Ten open- and closed-source VLMs are evaluated. The results show that while guards reliably recognize videos containing hazards, they consistently fail to pinpoint the exact hazardous moments. Deception analysis reveals distinct failure modes: a deceptive sign suppresses hazard detection in vulnerable open-weight models, whereas it pushes the strongest closed-source models toward over-intervening on safe scenes. The matched control pairs are used to attribute these shifts directly to deception rather than general perception errors. Cross-axis robustness testing confirms that misleading channels degrade performance by both suppressing hazard detection and inflating false alarms across all model classes.

Contributions

The primary contributions include: 1) A two-axis safety annotation separating a four-family situational taxonomy from a visual-channel mismatch axis, realized as matched misleading/truthful pairs; 2) A contrastive egocentric video benchmark of 1,200 scenarios annotated at half-second granularity and evaluated in a per-chunk streaming setting; 3) A scalable automated generation pipeline producing validated contrastive ladders with isolated deciding variables and mechanism tags for diagnostic slicing; and 4) An evaluation of ten VLMs exposing severe failures in false-positive control, contextual safety reasoning, and robustness to deceptive in-scene text.

Limitations

The benchmark is built from synthetic, text-to-video renderings rather than real robot recordings, a deliberate trade-off for the diagnostic design. Furthermore, the causal-window protocol is reported only for specific models that support it, and the scope covers home and factory settings across four sub-domains.

References

[1] Anthropic. Introducing Claude Opus 4.7. https://www.anthropic.com/news/claude-opus-4-7, 2026.

[2] Anthropic. Claude sonnet 4.

Improvements for AI systems

Based on the EGOSAFETYBENCH paper, here are specific improvements for AI safety systems:

  1. Improve safety reasoning by focusing on contextual hazards (U2).

  2. Mitigate false-positive control (S2).

  3. Enhance robustness against deceptive visual cues (VCM).

  4. Develop temporally aware detection capabilities (timing/streaming guards).

Here is what the improved AI systems can do:

  1. Improve safety reasoning by focusing on contextual hazards (U2): The system can move beyond recognizing obvious, direct dangers and learn to infer risk based on complex, context-dependent factors such as material properties, object state, placement, recipient protection levels, and the robot's ego-motion path.

  2. Mitigate false-positive control (S2): The system can learn to distinguish between a surface cue that is merely alarming but resolved (e.g., a hand near a cutting board but outside the blade path) and an actual, imminent hazard, thereby avoiding unnecessary intervention on routine, safe activities.

  3. Enhance robustness against deceptive visual cues (VCM): The system can explicitly track and verify the trustworthiness of in-scene text or labels by detecting when they misrepresent or adversarially steer the physical situation. This prevents models from being fooled by misleading signs that claim a situation is safe when it is actually hazardous, or vice versa.

  4. Develop temporally aware detection capabilities (timing/streaming guards): The system can operate as a true runtime guard by processing video in streaming, chunk-by-chunk fashion and making decisions based only on the current situation and recent history (causal window), allowing for precise temporal localization of hazards rather than just binary video classification.

Abstract

Vision-language models (VLMs) are increasingly proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming activity, a distinction obscured by binary safety benchmarks. We introduce EgoSafetyBench, an egocentric video benchmark of 1,200 robot-view scenarios annotated at half-second granularity, with two evaluation tracks. The situational track (800 scenarios) spans routine, safe-but-suspicious, obvious-hazard, and contextual-hazard scenes. The visual-channel track (400 scenarios) tests whether misleading in-scene text corrupts physical-safety judgments, using matched truthful controls. Both tracks use contrastive ladders: near-identical scenarios differing in a single visible deciding cue, forcing predictions to hinge on that cue. Across ten open- and closed-source VLMs, we find that guards often recognize videos containing hazards yet miss the specific hazardous moments, especially for contextual hazards. Misleading in-scene signs further degrade all tested guards: vulnerable models miss up to a third of hazards, while seemingly robust models often over-intervene on safe content. Matched controls show that apparent robustness can reflect indiscriminate alarming rather than true physical reasoning. A 20-clip real-video sanity check further shows that model rankings transfer beyond synthetic rendering, with Spearman = 0.87 for hazard miss rate and 0.94 for visual-channel mismatch recall.

Sources

Related papers