Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems
cs.RO, cs.AI, cs.CR
Submitted: 2026-09-02
Updated: 2026-09-02
DOI: 10.1007/978-3-032-33701-6_7
Code: https://github.com/doniobidov/silent_sabotage
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Composite Backdoor Attacks Against Large Language Models
- Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-based Decision-Making Systems
- Backdoor Attacks for In-Context Learning with Language Models
- Long-horizon Locomotion and Manipulation on a Quadrupedal Robot with Large Language Models
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic Trigger
- Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word Substitution
- Universal Jailbreak Backdoors from Poisoned Human Feedback
- BadGPT: Exploring Security Vulnerabilities of ChatGPT via Backdoor Attacks to InstructGPT
- Robot Collapse: Supply Chain Backdoor Attacks Against VLM-based Robotic Manipulation
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
- Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
- SafeEmbodAI: a Safety Framework for Mobile Robots in Embodied AI Systems
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving