Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations
cs.CL
Submitted: 2026-08-23
Updated: 2026-08-27
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
- Visibility into AI Agents
- Steering the distribution of agents in mean-field and cooperative games
- AI Research Considerations for Human Existential Safety (ARCHES)
- Open Problems in Cooperative AI
- Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents
- Group size effects and collective misalignment in LLM multi-agent systems
- Multi-Agent Risks from Advanced AI
- Conformity Generates Collective Misalignment in AI Agents Societies
- MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework
- Generative Agents: Interactive Simulacra of Human Behavior
- When Is Emergent Consensus Real? A Measured Coupling Gain and a Validity Diagnostic for LLM Agent Societies
- Improving Alignment and Robustness with Circuit Breakers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering