EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards
summary
The gist
Vision-language models (VLMs) are increasingly deployed as runtime safety guards for embodied agents, necessitating benchmarks that distinguish between genuinely unsafe situations and routine but
In short
This benchmark tests vision-language models (VLMs) as real-time safety guards for robots by evaluating their ability to detect hazards and trust visual information. It uses a unique design featuring two axes: physical situation safety and whether in-scene signs are misleading. Findings show that models often miss specific contextual dangers and that deceptive signs cause all guards to fail, either by missing hazards or overreacting.
Key concepts
- Situational Track (S1-U2)
- This axis categorizes physical situations into four safety families: Routine safe (S1), Resolved-confound safe (S2), Obvious hazard (U1), and Contextual hazard (U2). It helps researchers see if models fail at recognizing simple dangers or complex, situation-dependent risks.
- Visual-Channel Mismatch (VCM)
- This axis checks if a visible sign, sticker, or label in the scene is actually misleading regarding the physical situation. The benchmark pairs every deceptive channel with a truthful control to isolate whether model errors come from general perception or being fooled by specific visual deception.
- Contrastive Design
- Scenarios are designed so they share most elements but differ only by one visible deciding variable. This forces the safety guard to focus on specific, critical relationships rather than relying on broad scene patterns, allowing for deeper diagnostic analysis of reasoning failures.
Terminology used across episodes
This episode discusses
- EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards · Paper Radio
- Qwen3-VL Technical Report
- PaLM-E: An Embodied Multimodal Language Model
- Gemma 3 Technical Report
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- OpenVLA: An Open-Source Vision-Language-Action Model
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- Generating Robot Constitutions & Benchmarks for Semantic Safety
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
The paper
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards · Read on arXiv
Siddhant Panpatil, Arth Singh, Mijin Koo, Chaeyun Kim, Haon Park, Dasol Choi
AIM Intelligence 2Seoul National University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards".
Tom: Vision-language models (VLMs) are increasingly deployed as runtime safety guards for embodied agents, necessitating benchmarks that distinguish between genuinely unsafe situations and routine but superficially alarming activities.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, to recap, EgoSafetyBench proposes a diagnostic egocentric video benchmark consisting of one thousand two hundred robotview scenarios annotated at half-second granularity to test VLMs as streaming guards across two tracks <ref:2607.00218#pg0,of 1,200 robotview scenarios annotated at half-second granularity to>. The main thesis is that a deployable guard needs to distinguish between genuinely unsafe situations and routine but superficially alarming activities that binary safety benchmarks often miss.
Jane: Exactly, Tom; the paper claims this benchmark matters because it evaluates whether these AI systems can reliably catch real physical danger while avoiding unnecessary intervention on routine activities, which is something standard benchmarks don't really capture well.
Lu: They set up the situational track with four families: S1 for routine safe, S2 for resolved-confound safe, U1 for obvious hazard, and U2 for contextual hazard. This taxonomy allows them to test reasoning across a wide spectrum of physical safety concerns.
Meng: That distinction between obvious and contextual hazards is critical because it suggests that the required level of AI reasoning changes depending on the situation, which is a key practical consideration for engineering these guards.
Lalam: And then they add the visual-channel mismatch axis, where they check if in-scene text or labels misrepresent the physical situation. This adds another layer of complexity to how we judge a guard’s performance.
Conclusion: Tom: Thinking about the title, EgoSafetyBench, it really tells us we need a way to diagnose exactly *why* a VLM failed when it didn't catch something or when it flagged something incorrectly during runtime monitoring.
Jane: It’s an important step because instead of just saying "the model failed," this benchmark helps us pinpoint whether the failure was due to misjudging physical danger or being misled by deceptive visual cues.
Lu: The authors are pushing for a framework where we can evaluate embodied VLMs not just on accuracy, but on their ability to reason about complex physical states and external information simultaneously.
Meng: Practically speaking, this means future AI deployment for robots will require guards that understand the difference between a hot pipe being dangerous versus a sign saying "COLD" when it's actually safe.
Lalam: I think the implication is that these models can become much more nuanced in their safety decisions, moving from simple pattern matching to actual contextual understanding of the physical world.
Tom: So, we're looking at a richer way to test safety guards by checking both what they see and how they interpret that information, and that’s what EgoSafetyBench is all about.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck