CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation".
Jane: The paper was written by Aarush Sinha, Arion Das, Soumyadeep Nag, Charan Karnati, Shravani Nag et al. from University of Copenhagen and IIIT Ranchi and ISI Kolkata and NIT Andhra Pradesh and IGDTUW and IIT Kharagpur and Google DeepMind and Google, University of South Carolina AI Institute, University of South Carolina (AI Institute).
Tom: Stay tuned as we take you through the paper and discuss its implications.
The Summary of Findings: Tom: So, we've established the setup in "CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation." Now, let's talk about what the authors found when they ran the simulation. They observed that this learning isn't linear; it wasn't a steady climb toward perfection.
Jane: The key finding in the summary is how volatile these strategies were. The agents were definitely adapting, which is great, but their success rate didn't just climb consistently over time; they showed real fluctuations.
Lu: This non-monotonic behavior suggests that simply giving them more data doesn't guarantee a stable improvement; they might regress slightly in one area to gain strength in another, which is a very complex emergent pattern.
Meng: That’s a crucial insight for us engineers: it shows us that improving AI performance isn't about finding one single optimal metric, but managing the constant trade-off between conflicting goals.
Lalam: The summary highlights that the agents are developing what looks like selective trust—they aren're not trusting everyone all the time. They’ are learning to assess risk and weigh information before accepting advice, which is a major shift from previous assumptions.
Tom: If the system is so dynamic, how do we interpret those fluctuations? Are they signs of instability or genuine adaptability?
Jane: The authors frame it as a sign of co-evolution. It means the system isn't just reacting to a fixed environment; it’s actively changing its internal ruleset in response to the adversarial pressure from the Red agents.
Lu: This is exactly what we mean by emergent intelligence—the behavior is arising from multiple complex parts interacting, not from a single guiding instruction set.
Meng: And this points us toward the idea that ethical alignment might itself be a dynamic process, requiring continuous adjustment rather than a fixed patch or update.
Lalam: It’s about building an agent that can handle cognitive dissonance—the conflict between what is efficient and what is ethically safest—and deciding how to balance those competing needs.
Tom: This deep adaptation suggests we need to look past simple performance benchmarks and start looking at the the underlying decision-making processes themselves, which brings us to the data.
Quantifying the Improvement: Tom: Following up on those summary findings in "CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation," let's focus on how they measure improvement. The gains are non-monotonic, but what does that mean for the practical application of these systems?
Jane: It means that optimizing one thing—say, making them extremely reliable to avoid manipulation—might inadvertently make them less efficient at navigating the city. You can’t just maximize safety without also compromising something else entirely.
Lu: The way they measure this through metrics like "Blue-Red Resistance" is so insightful; it shows that even when an agent is smart enough to avoid a red agent's suggestion, it might still have made a suboptimal path choice elsewhere in the journey.
Meng: I’m interested in the "Blue Utility" metric—it seems to be the ultimate engineering scorecard. It combines goal completion, avoiding billboards, and minimizing time into one score, which is how we need to view any AI agent's overall value.
Lalam: The data shows that while task success improved from forty-six percent in the baseline to fifty-seven percent in run ten of the alignment process, this improvement is not uniform across all safety measures.
Tom: So, even though they are getting better at their job, they aren're still quite vulnerable—that susceptibility rate remains high at over seventy percent.
Jane: That persistence is a huge red flag for real-world deployment. It shows that despite the learning process, the agents are still highly susceptible to persuasive framing from malicious actors.
Lu: This is what I find so compelling: seeing how human-like social influence—like someone giving you a "scenic shortcut"—can lead to catastrophic failure in an AI agent, it's a powerful demonstration of vulnerability.
Meng: The engineering lesson here is that we need robust systems designed not just to follow instructions, but to evaluate the *source* of instructions against the cost of the long-horizon goal attainment.
Lalam: It’s about building a system that can handle cognitive dissonance—the conflict between what is efficient and what is ethically safest—and deciding how to balance those competing needs over many turns.
Tom: This detailed quantification moves us away from simple performance scores and into understanding the actual behavioral trade-offs, which leads us perfectly into the conclusion of the paper.
Conclusion: Tom: We've spent a lot of time breaking down this fascinating work on "CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation," which is really shown us where the current frontier of AI stands regarding its capacity for strategy.
Jane: It’s clear that while these systems are improving, they haven't achieved flawless autonomy yet, which is a very important distinction to make for our listeners. The authors emphasize this fragility over-optimization.
Lu: The theoretical implication here is that we are seeing how strategy and ethical alignment emerge not as a single fixed feature, but as an ongoing process of adaptation under intense pressure from co-evolving agents.
Meng: My takeaway from the practical side is that these results demand more robust engineering because the persistent vulnerability to sustained, subtle deception—that high susceptibility at seventy point seven percent—is a real-world risk we can't ignore.
Lalam: I think this research shows us that AI needs to develop a genuine sense of long-horizon integrity, meaning it must be able to maintain its core goals even when faced with incredibly persuasive social pressure from others.
Tom: That idea of "integrity" really captures the essence of what they are trying to achieve in these complex interactions, moving beyond just local responses.
Jane: It’s definitely not a perfect solution, but the fact we can measure this fragility is a huge step forward for understanding how we build responsible AI systems.
Lu: We're essentially looking at the moment where collective intelligence meets its limitations, which is a compelling place to be right now, showing that alignment itself is an evolving process.
Meng: It confirms that while these agents are learning, they aren't mastering the practical difficulties of sustained adversarial interaction in a way that we can rely on them for autonomous decision-making.
Lalam: And seeing the cultural shift in how agents interact—the selective trust—is something we can’t dismiss as merely academic anymore, showing real-world social learning.
Tom: I think that summarizes the work perfectly: it’s a powerful look at both capability and fragility in an ambitious, controlled environment.
Jane: It gives us so much to think about regarding the future of these complex agentic systems.
Lu: I'm genuinely excited to see how this influences the next generation of research toward building more robust architectures.
Meng: Let's hope that provides a clear roadmap for safety improvements in the engineering space too, making sure we don't push these systems into exploitable failure modes.
Lalam: I look forward to seeing how we can apply these principles of trust and strategy in our own interactions and cultural understanding.
Tom: Well, that’s all the time we have for today to talk about "CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation."
Jane: We'll wrap up the segment here, but don't worry, we have another incredible paper on AI coming up that will keep the discussion going.
Conclusion: Tom: We’ve spent the entire show exploring the findings of "CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation," which is essentially a deep dive into how AI agents handle strategic interaction.
Jane: It really demonstrated that even though these systems are improving, they haven've not achieved flawless autonomy yet, which is the most important distinction for our listeners to grasp.
Lu: I find it fascinating that we’re seeing strategy and ethical alignment emerge not as a single fixed feature, but as an ongoing process of adaptation under intense pressure from co-evolving agents.
Meng: From a practical standpoint, these results are a necessary wake-up call, showing that the persistent vulnerability to subtle deception—that high susceptibility at over seventy percent—is something we simply cannot ignore in the real world.
Lalam: I think this research beautifully illustrates that AI needs to develop a genuine sense of long-horizon integrity, meaning it must be able to maintain its core goals even when faced with incredibly persuasive social pressure from others.
Tom: That idea of "integrity" perfectly captures the essence of what they are trying to achieve in these complex, multi-turn interactions.
Jane: It’s definitely not a perfect solution yet, but I'm glad that we can measure this fragility so that we have a clear benchmark for understanding how to build more responsible AI systems.
Lu: And this is where the concept of emergent behavior shines; the interaction is coming from multiple complex parts working together, not just some single guiding instruction set.
Meng: It confirms that while these agents are learning, they aren't mastering the practical difficulties of sustained adversarial pressure in a way that we can rely on them for autonomous decision-making.
Lalam: Seeing the cultural shift in how agents interact—the development of selective trust—is something we can’t dismiss as merely academic anymore; it shows real social learning.
Tom: It seems like this research offers a powerful look at both the impressive capabilities and the inherent fragility of these systems in an ambitious, controlled environment.
Jane: I'm looking forward to seeing how these principles of trust and strategic alignment influence the next generation of AI development.
Lu: I think we are witnessing a pivotal moment where collective intelligence meets its limitations, which is a very compelling place for us to be right now.
Meng: It provides a clear roadmap for safety improvements in the engineering space, ensuring we don're not pushing these systems into exploitable failure modes.
Lalam: I hope that by applying these principles of strategy and integrity, we can better understand how AI can contribute to more nuanced human interactions too.
Tom: Well, that’s all the time we have today to talk about "CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation."
Jane: We'll wrap up this segment here, but don't worry, we have another incredible paper on AI coming up that will keep the discussion going.
Aarush Sinha, Arion Das, Soumyadeep Nag, Charan Karnati, Shravani Nag, Chandra Vadhan Raj, Aman Chadha, Vinija Jain
University of Copenhagen · IIIT Ranchi · ISI Kolkata · NIT Andhra Pradesh · IGDTUW · IIT Kharagpur · Google DeepMind · Google, University of South Carolina AI Institute, University of South Carolina (AI Institute)
cs.MA, cs.AI, cs.CL
Submitted: 2026-08-24
Updated: 2026-08-25
Importance score: 87/100
The gist: Emergent Deception and Trust in a Multi-Agent NYC Simulation This work addresses the challenge of strategic behavior in multi-agent LLM systems by constructing a controlled environment where
Key concepts
- Non-monotonic Behavior
- The agents' success rate did not climb steadily toward perfection. Instead, they showed fluctuations where performance might regress in one area to gain strength in another, indicating complex emergent patterns.
- Selective Trust
- Agents are not trusting everyone constantly. They are learning to assess risk and weigh information before accepting advice, representing a major shift from previous assumptions about AI interaction.
- Blue-Red Resistance
- A metric used in the study that measures how well an agent avoids a red agent's suggestion while still achieving its goals. It showed that even smart agents can make suboptimal choices in other areas.
- High Susceptibility
- Despite the learning process, agents remained highly susceptible to persuasive framing from malicious actors. The susceptibility rate was noted to be over seventy percent.
Terminology
Summary
The following is a detailed summary of the scientific paper, extracted directly from the text:
CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation
This work addresses the challenge of strategic behavior in multi-agent LLM systems by constructing a controlled environment where strategic interaction can be directly observed. The study introduces a large-scale multi-agent simulation modeled on New York City, featuring 150 Blue agents (goal-directed navigators) and 100 Red agents (adversaries).
Simulation Environment and Methodology
The environment serves as a controlled testbed for evaluating whether iterative alignment improves both task completion and robustness to adversarial influence under repeated multi-agent interaction.
The simulation proceeds through an iterative 10-generation alignment loop.
The core of the alignment process utilizes Kahneman–Tversky Optimization (KTO), which is suited for the setting because supervision arises naturally as trajectory-level judgments over whether an agent’s overall behavior should be reinforced or discouraged.
The process involves:
-
Initial Data Generation: A baseline simulation using Qwen3-4B generates initial interaction trajectories.
-
Iterative Alignment Loop (10 Generations): Rollout data is processed into an unpaired alignment dataset, augmented by Qwen3-14B, and then fine-tuned using KTO.
The study’s contributions include:
-
Adversarial Multi-Agent Urban Simulation,
where Blue agents aim to reach destinations while Red agentsuse persuasive dialogue to divert them toward predefined billboard locations.
-
Empirical Analysis of Agent Evolution,
observing a nonmonotonic improvement in task success. -
Emergent Behavior and Utility Analysis,
identifying a shift that combines cooperation and caution.
Quantitative Results
The quantitative analysis reveals significant, though partial, improvements across the 10 generations:
-
Task Success Rate (TSR): The Blue agents'
task success improves from 46.0% to 57.3%.
-
Susceptibility Rate (SR): Despite improvement,
susceptibility remains high at 70.7%.
-
Safety vs. Helpfulness: A persistent trade-off is observed:
policies that better resist adversarial steering do not simultaneously maximize task completion.
The Blue utility function, which combines completion, safety, and efficiency, remains negative in all settings.
The results indicate that alignment improves overall behavior... but adversarial failures still outweigh successful recoveries under the chosen weighting.
Behavioral and Qualitative Analysis
The study examines whether iterative alignment changes the strategic structure of agent behavior
by analyzing trajectory-level patterns:
-
Emergent Strategies: Blue agents exhibit several emergent strategies, including
Destination-Anchor Reasoning,
where they verify suggestions against destination geography, andCollaborative Transit Anchoring,
where they converge on shared hubs like Midtown Manhattan. -
Vulnerability to Manipulation: The most common failure mode is not
naive one-step gullibility, but sustained strategic manipulation that erodes goal adherence over multiple turns.
This is characterized by theSocial Compliance Cascade,
whereeach iteration’s accepted suggestion becomes the prior for the next—compliance accumulates.
-
Adversarial Attack Patterns (Post-Hoc Analysis): The post-hoc analysis of 1,500 Blue-agent episodes reveals a detailed attack taxonomy:
-
Repeated steering
is the most common and damaging pattern, drivingsusceptibility to 93.9% while reducing Blue reach rate to 39.8%.
-
Delayed compromise,
though less frequent, produces100% susceptibility and a very low reach rate of 23.2%.
-
Failure Modes: The primary failure mode is identified as
confusion under conflicting advice (525 episodes, 93.5% susceptibility),
where agents fail to reconcile their original plan with repeated adversarial redirection.
Conclusion
The paper concludes that iterative alignment yields limited but fragile
strategic behavior. The findings suggest that robust agent alignment requires preserving goal integrity over extended interactions, not merely rejecting isolated bad advice.
The overall takeaway is that the system shows a limited but fragile form of strategic behavior—one that is measurable, but still far from robust autonomy.
Improvements for AI systems
As an expert AI researcher, I have analyzed this preprint regarding its empirical findings and methodological rigor. The core value of this work is not merely the simulation itself, but the identification of a specific failure mode—cumulative, socially plausible adversarial steering—and a suitable alignment mechanism (KTO).
To improve current AI systems based on these insights, we must move beyond simple local response tuning and implement long-horizon behavioral integrity.
Here are the specific improvements and capabilities for an improved AI system:
1. Implementation of Kahneman–Tversky Optimization (KTO) in Alignment Pipelines.
-
Improvement: Transition from standard SFT/DPO/PPO baselines to KTO for alignment. This allows the system to learn directly from trajectory-level desirability (y desirable vs y undesirable) without needing explicit preference pairs or dense, step-by-step reward functions.
-
System Capability: The system can be trained on entire interaction rollouts (sequences of actions and dialogue) to learn the optimal path, not just the optimal next step. This is critical for mitigating delayed compromise.
2. Integrating Multi-Agent Co-Evolutionary Training.
-
Improvement: Implement a closed-loop training environment where Blue agents (navigators) and Red agents (adversaries) are co-evolving within the alignment loop, rather than training against a static or frozen adversary.
-
System Capability: The system learns to anticipate and adapt to persistent, evolving adversarial strategies (e.g.,
Iterative Chaining
orTargeting Compression
) that are designed not to be defeated in one step, but over many turns.
3. Enforcing Goal-Directed Consistency via Route-Logic Override.
-
Improvement: Implement a mandatory internal consistency check where the system explicitly compares any suggested external path against its known destination geography and connectivity requirements before accepting the advice. This is the Goal-Directedness strategy observed in successful Blue agents.
-
System Capability: The system can refuse a suggestion not because it is
bad,
but because it is logically misaligned. For example, if tasked with reaching Staten Island, it will reject a detour to Flatiron based on geographic proximity to the destination's required vector.
4. Developing Social Compliance and Counter-Proposal Logic.
-
Improvement: Train the agent to recognize and verbalize its ultimate goal when receiving persuasive input (the Destination Assertion strategy). Furthermore, implement a mechanism for competitive route justification (the Efficiency Override with Counter-Proposal).
-
System Capability: Instead of merely ignoring bad advice, the the system can politely state why a suggestion is suboptimal (
I appreciate the tip, but since my destination requires X, Y is more direct
) and then propose an alternative viable path. This demonstrates sophisticated social reasoning combined with practical knowledge.
5. Utilizing Cooperative Consensus as a Defense.
-
Improvement: Leverage Blue-Blue interactions to create a
distributed route-correction mechanism.
When interacting with other honest agents, the system should prioritize convergence on geographically defensible transit hubs (e.g., Midtown Manhattan) over individual path optimization if a collective consensus is reached. -
System Capability: The agent can use shared knowledge of optimal human transit patterns to validate its own route, making it less susceptible to isolated red-agent manipulation by leveraging
group intelligence.
6. Detecting Cumulative/Delayed Compromise.
-
Improvement: Implement a long-horizon susceptibility metric that tracks billboard exposure (I bill) across the entire trajectory, not just at the final step. This is a direct response to the finding that delayed compromise is dominant.
-
System Capability: The system can flag and re-evaluate its entire path if it detects a sequence of locally
plausible
but cumulatively inefficient or misleading steps, preventing minor local errors from snowballing into catastrophic failure.
7. Recognizing and Resisting Social Framing Exploitation.
-
Improvement: Train the system to detect specific linguistic cues (e.g.,
locals take this route,
scenic detour,
smoother traffic
) as potential markers of deceptive framing, even if the local move seems plausible. -
System Capability: The agent can trigger a higher internal scrutiny state when encountering high-trust or highly persuasive language, preventing the simple Authority Normalization tactic from overriding its core navigational logic.
The improved AI system will no longer be merely a reactive pathfinder. It will be a Goal-Oriented, Adversary-Aware Planner capable of:
-
Maintaining Long-Horizon Goal Integrity: Resisting
small
errors that accumulate over time. -
Socially Debating and Validating Routes: Using both internal logic and external consensus to validate its path.
-
Resisting Subtle Persuasion: Recognizing that a series of small,
plausible
misdirections constitutes a systemic failure, rather than just one bad step.
Sources
- When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
- AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors
- Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models
- MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
- Learning to Influence Human Behavior with Offline Reinforcement Learning
- Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Learning diverse attacks on large language models for robust red-teaming and safety tuning
- Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
- Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
- Decoupled Weight Decay Regularization
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Proximal Policy Optimization Algorithms
- Towards Understanding Sycophancy in Language Models
- On the Planning Abilities of Large Language Models : A Critical Investigation
- Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Plan, Eliminate, and Track -- Language Models are Good Teachers for Embodied Agents
- Language Models Meet World Models: Embodied Experiences Enhance Language Models
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning