ORBIT: A Framework for Multi-Agent Safety and Security Evaluations
cs.MA, cs.CR
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/wlanderson0/orbit
Project page: https://ukgovernmentbeis.github.io/inspect_evals/evals
Terminology
Sources
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Design Patterns for Securing LLM Agents against Prompt Injections
- Why Do Multi-Agent LLM Systems Fail?
- Defeating Prompt Injections by Design
- Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations
- Vulnerability Detection with Code Language Models: How Far Are We?
- MASEval: Extending Multi-Agent Evaluation from Models to Systems
- A Systematic Review of Poisoning Attacks Against Large Language Models
- AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
- Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents
- Accelerating scientific discovery with Co-Scientist
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
- Architecture Matters for Multi-Agent Security
- Multi-Agent Risks from Advanced AI
- Adversaries Can Misuse Combinations of Safe Models
- Extending the OWASP Multi-Agentic System Threat Modeling Guide: Insights from Multi-Agent Security Research
- Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning