Trust the Brand, Lose Control: How Identity Hijacks LLM Agent Orchestration
cs.CR, cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/anomalyco/opencode
Terminology
Sources
- Exposing Weak Links in Multi-Agent Systems under Adversarial Prompting
- The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models
- RouterBench: A Benchmark for Multi-LLM Routing System
- Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
- SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
- In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations
- LLMs Corrupt Your Documents When You Delegate
- AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments
- Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- Auditing Agent Harness Safety
- LDP: An Identity-Aware Protocol for Multi-Agent LLM Systems
- AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration
- Multi-Agent Collaboration Mechanisms: A Survey of LLMs
- OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
- ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
- Agent-as-a-Router: Agentic Model Routing for Coding Tasks
- MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs