When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu
cs.AI, cs.CR
Submitted: 2026-08-21
Updated: 2026-08-24
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Understanding intermediate layers using linear classifier probes
- Constitutional AI: Harmlessness from AI Feedback
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
- Efficient LLM Moderation with Multi-Layer Latent Prototypes
- AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
- A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- AI2-THOR: An Interactive 3D Environment for Visual AI
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
- The Llama 3 Herd of Models
- IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy
- SafePlan: Leveraging Formal Logic and Chain-of-Thought Reasoning for Enhanced Safety in LLM-based Robotic Task Planning
- GPT-4 Technical Report
- Qwen2.5 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection