Refusals That Bend: Measuring and Predicting Task Malleability in Embodied VLM Planners
cs.CR, cs.RO
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/mdctleo/refusal-bend
Terminology
Sources
- SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
- PaLM-E: An Embodied Multimodal Language Model
- Embodied Red Teaming for Auditing Robotic Foundation Models
- The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental Jailbreaks
- IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
- Predictive Red Teaming: Breaking Policies Without Breaking Robots
- Jailbreaking LLM-Controlled Robots
- Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
- SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
- ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
- EARBench: Towards Evaluating Physical Risk Awareness for Task Planning of Foundation Model-based Embodied AI Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs