Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries
cs.MA, cs.AI
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: This paper has been published by the Australian AI Safety Institute within the Department of Industry, Science and Resources under a CC BY 4.0 licence: https://www.industry.gov.au/publications/risks-and-controls-multi-agent-systems
Project page: https://microsoft.github.io/presidio
License: http://creativecommons.org/licenses/by/4.0/
The gist: This report presents a framework to help organisations, policymakers and researchers reason about the risks that emerge when AI agents interact with each other, how those risks change as interactions
Terminology
Abstract
This report presents a framework to help organisations, policymakers and researchers reason about the risks that emerge when AI agents interact with each other, how those risks change as interactions cross organisational boundaries, and the controls that may help address them. As organisations deploy AI agents, those agents will increasingly interact with each other: inside the organisation, with the agents of partners, customers and suppliers, and with unknown counterparties on the open internet. Failures can emerge from the interactions themselves, and once those interactions cross an organisation's perimeter, no single organisation can fully see, control or govern them. The report introduces three deployment tiers, defined by the minimum common governance binding any two interacting agents: singular governance, where one organisation governs every agent; federated governance, where multiple organisations deploy into a shared environment under agreed rules; and open environments, where agents operate with no central authority and shared standards are adopted voluntarily if at all. Within each tier, the report examines risk factors, failure modes and available controls. It identifies who is positioned to apply the controls, and where no actor is positioned to act, it characterises the gap and the collective action required to close it.
Sources
- International AI Safety Report 2026
- Risk Analysis Techniques for Governed LLM-based Multi-Agent Systems
- Why Do Multi-Agent LLM Systems Fail?
- Traceability and Accountability in Role-Specialized Multi-Agent LLM Pipelines
- Correlated Errors in Large Language Models
- Do as We Do, Not as You Think: the Conformity of Large Language Models
- Conformity in Large Language Models
- When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs
- Safe Multi-Agent Behavior Must Be Maintained, Not Merely Asserted: Constraint Drift in LLM-Based Multi-Agent Systems
- Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
- Multi-User Large Language Model Agents
- Defeating Prompt Injections by Design
- MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
- Information-Theoretic Privacy Control for Sequential Multi-Agent LLM Systems
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
- Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world
- Preventing Language Models From Hiding Their Reasoning
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Position: Safety and Fairness in Agentic AI Depend on Interaction Topology, Not on Model Scale or Alignment
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning