The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
cs.CL, cs.CY
Submitted: 2025-09-22
Updated: 2026-09-10
Comments: Added more studies in our systematic audit (350 papers; 576 simulations)
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- Herd Behavior: Investigating Peer Influence in LLM-based Multi-Agent Systems
- Large Language Models can Achieve Social Balance
- LLM Social Simulations Are a Promising Research Method
- Emergence of Scale-Free Networks in Social Interactions among Large Language Models
- Emergent social conventions and collective bias in LLM populations
- SycEval: Evaluating LLM Sycophancy
- The Power of Stories: Narrative Priming Shapes How LLM Agents Collaborate and Compete
- Large Language Model Driven Agents for Simulating Echo Chamber Formation
- GPT-4o System Card
- Diversity of Thought Elicits Stronger Reasoning Capabilities in Multi-Agent Debate Frameworks
- Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
- Evolution of Social Norms in LLM Agents using Natural Language
- Exploring Social Desirability Response Bias in Large Language Models: Evidence from GPT-4 Simulations
- Simulating Rumor Spreading in Social Networks using LLM Agents
- Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges
- Large Language Model-driven Multi-Agent Simulation for News Diffusion Under Different Network Structures
- Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs
- Large Language Models Often Know When They Are Being Evaluated
- Probing and Steering Evaluation Awareness of Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering