Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes

arXiv:2608.11207 · cs.AI · Submitted 2026-04-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes".

Jane: The paper was written by Alexander Liss, Nicholas Desmond and Santiago Gil Gallego from Georgia Institute of Technology and Huge Inc..

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's been making the rounds, and I gotta say, the title alone got me hooked: "Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes." Jane, when you first saw that title, what went through your head?

Jane: Honestly, Tom, I had to read it twice. But once I did, I realized it's asking a question that's been bugging me for a while. We keep putting multiple AI agents in rooms together and expecting them to play nice, but nobody's really figured out how to make that work reliably. This paper says, hey, maybe we need a traffic controller.

Tom: A traffic controller, I like that. And that's exactly what they built. They call it the Experience Orchestrator, or EO for short. It's like a governance layer that sits above two AI agents and steers their conversation toward a goal.

Jane: Right, and the setup is pretty clever. You've got a financial services website, and there's a chatbot trying to get a visitor to schedule a meeting with a financial advisor. But the visitor isn't just going to roll over — it's another AI agent with its own persona, its own skepticism, its own reasons to resist.

Tom: So it's like two AIs having a conversation, and neither one knows what the other is trying to do. The chatbot wants a conversion, the visitor wants to maintain its resistance. And without any outside help, what happens?

Jane: They just drift. The paper calls it convergence toward agreement, but it's not real agreement. The visitor might say yes just to end the conversation, or the chatbot gives up and they both settle into a polite but useless exchange.

Tom: And that's the core problem the paper is tackling. When you have two AI agents with opposing goals, and there's no shared reward function — no mathematical target they're both optimizing for — they don't compete, they collapse. They end up in a state that satisfies neither of them.

Jane: Exactly. And that's where the governance layer comes in. It's not changing the AI models themselves. It's not retraining them. It's observing the conversation from the outside and applying corrective force, like a thermostat keeping a room at the right temperature.

Tom: A thermostat, I love that analogy. And the results are pretty dramatic. We're talking about a thirty-two percentage point lift in getting visitors to actually contact an advisor, compared to a naive chatbot that's just given a system prompt and left to its own devices.

Jane: That's the headline number, but what's even more interesting to me is what they found about why it works. The governance policy — the choices the system makes about how to steer the conversation — accounts for ninety-seven percent of the variation in outcomes. Not the environment, not the visitor's starting attitude, not the persona. The steering.

Tom: So the way you govern the conversation matters almost everything, and the starting conditions barely matter at all. That's a bold claim, and we're gonna dig into how they actually pulled that off. But first, let's talk about who wrote this thing.

Jane: Alexander Liss from Georgia Tech, along with Nicholas Desmond and Santiago Gil Gallego from Huge Inc. And you can tell it's a collaboration between academia and industry, because it's got both the theoretical framework and the practical engineering.

Tom: And that combination is exactly what we need right now. We're about to see a flood of multi-agent AI systems in the real world, and this paper is one of the first serious attempts to figure out how to keep them on track. Stick around, because next we're gonna break down what they actually built and how the simulation works.

Summary: Tom: Welcome back. We're still on "Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes," and Jane, I want to get into the meat of it. How did they actually test this thing?

Jane: So they built a massive simulation. We're talking over sixty thousand simulated visitor sessions on a financial services website. Each session has a visitor agent with a specific persona — there are six of them, things like the fee hawk who's analytically skeptical, or the stranded saver who's anxious about retirement.

Tom: And each persona has its own way of responding to different persuasion strategies. The fee hawk wants hard numbers, the legacy loyalist wants trust and continuity, the grieving proxy is emotionally vulnerable and needs empathy, not a sales pitch.

Jane: Right. And the visitor navigates through a simulated website, visiting pages based on real web analytics data. Then they hit this decision boundary page, and that's when the chatbot appears and tries to guide them toward contacting an advisor.

Tom: But here's the twist — the chatbot doesn't just wing it. It's got this governance layer watching the whole conversation, and that layer is made of three parts. Can you walk us through them?

Jane: Sure. First, there's a PID controller — that's a classic engineering feedback mechanism. It watches the visitor's resistance level, which is basically how closed off they are to the idea of talking to an advisor. If resistance is high, the controller pushes harder. If it's stuck, it changes strategy.

Tom: And the second part?

Jane: A belief tracker. The chatbot can't actually read the visitor's mind, so it maintains a probability distribution over what the visitor might really want — are they just browsing, comparing options, ready to buy, or looking for support? As the conversation goes on, it updates that belief based on what the visitor says.

Tom: And the third piece is the contextual bandit. That's the decision-maker. At each turn, it picks one of four content strategies — direct contact pitch, trust-building guidance, math and numbers, or friction reduction. And it picks based on everything it knows about the visitor so far.

Tom: So it's like the bandit is the brain, the PID is the hand on the wheel, and the belief tracker is the eyes. And they all work together to keep the conversation moving toward a good outcome.

Jane: Exactly. And the results are striking. The full system gets seventy-eight percent of high-intent visitors to actually contact an advisor, compared to forty-six percent for a naive chatbot that's just given a prompt and told to do its best.

Tom: But here's what really blew my mind. They ran this factorial experiment across eight different friction models — basically eight different assumptions about how visitor resistance evolves — and six personas. And they found that the choice of governance policy explains ninety-seven percent of the variance in outcomes. The friction model only explains three percent.

Jane: That's the big finding. It means the environment barely matters compared to the quality of the steering. You could have the most realistic simulation in the world, but if your governance policy is bad, you're gonna get bad results. And if your governance policy is good, it can overcome a lot of environmental noise.

Tom: So it's not about the starting conditions, it's about the control. And that's a really hopeful message for people trying to build these systems in the real world, because you can't control your environment, but you can control your governance.

Jane: And they even showed what happens when governance is missing. The conversations just drift into these failure modes — sycophantic collapse where the visitor agrees without meaning it, stagnation where nothing moves, or the chatbot getting fixated on one strategy and ignoring everything the visitor says.

Tom: Those failure modes sound familiar, honestly. I feel like I've had customer service chats that hit all three. But the fact that they can name them and detect them means they can correct them.

Jane: And that's the real contribution here. It's not just a better chatbot — it's a framework for thinking about how to govern any multi-agent system. And that's what we're gonna dig into next, because the implications go way beyond financial services.

Tom: You're reading my mind. Let's talk about what this means for the rest of the AI world after the break.

Improvements: Tom: We're back with "Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes," and I want to bring in our regulars now, because this paper has implications that go way beyond chatbots on financial websites. Lu, you've been quiet — what's jumping out at you?

Lu: Tom, what excites me is the generalization. The authors frame this as "adversarial-adjacent" multi-agent dynamics — two agents with opposed goals, but where cooperation is possible if you steer correctly. That's not just financial services. That's healthcare consultations, enterprise software sales, HR recruiting, technical support. Anywhere you have one agent trying to guide and another agent maintaining resistance.

Jane: And the key insight is that you don't need to retrain the agents. You don't need them to share a reward function. You just need a better governor sitting on top. That's huge for practical deployment, because retraining is expensive and risky.

Meng: But let me push back a little, because I'm the engineer who has to actually run this thing. The PID controller is calibrated against an LLM that reliably self-reports its resistance score on a structured scale. Real humans don't do that. You can't just ask a visitor "on a scale of one to five, how resistant are you to talking to an advisor?" and expect a useful answer.

Tom: That's a fair point, Meng. The paper acknowledges this limitation pretty directly. The visitor agent is another language model, not a human. Real visitors are far more unpredictable — they don't maintain consistent personas, they respond to subtext, they might just get bored and leave.

Meng: Right. And the PID gains — the specific tuning parameters — were chosen based on how an LLM behaves. A human might escalate, disengage, or do something that falls entirely outside the six archetypes they modeled. So the thirty-two-point lift is real, but it's measured in a sandbox.

Jane: But Meng, isn't that true of most simulation research? You have to start somewhere, and the fact that they're being upfront about the limitation is actually a good sign. They're not claiming this is production-ready. They're saying this is a proof of concept.

Lu: And I'd add that the variance decomposition result — ninety-seven percent policy, three percent environment — is actually more robust to that criticism than the absolute numbers. Even if the friction models don't perfectly match human behavior, the fact that governance dominates environmental factors across eight different models suggests the finding is structural, not an artifact of one simulation choice.

Meng: Okay, that's fair. And I'll admit, the architecture is elegant. The contextual bandit picks the strategy, the PID shapes the response, the belief tracker updates the model of the visitor. It's a clean separation of concerns. And they even handle the sparse reward problem with a shaped reward that gives feedback at every turn instead of just at the end.

Tom: So what's the path to making this work with real humans?

Meng: The authors propose live A/B testing, which is the obvious next step. And they mention something interesting — a companion paper that uses attention fine-tuning to derive a reward signal from the model's own internal dynamics, instead of relying on a PID heuristic. That could make the governance more adaptive to individual visitors.

Lu: And I'd push even further. Imagine extending this with POMCP lookahead planning — simulating several turns ahead before choosing a strategy. That would let the system anticipate resistance escalation before it happens, instead of reacting to it after the fact.

Jane: So the improvements are already mapped out. Live validation, better reward signals, longer planning horizons, and testing across different domains. This paper is really a foundation, not a finished building.

Tom: And that's the exciting part for me. We're watching a new field take its first steps. Before we wrap up, I want to bring in Lalam to get a broader cultural perspective on what this means.

Conclusion: Tom: We're in the final stretch on "Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes," and I want to hear from Lalam before we say goodbye to this paper.

Lalam: Tom, what I find most impactful is the shift in mindset this paper enables. For a long time, we've thought about AI alignment as something you bake into the model during training. This paper suggests you can also govern behavior from the outside, at runtime, without touching the weights. That's a different philosophy entirely.

Jane: It's like the difference between raising a child to always make good choices versus having a responsible adult in the room who can gently redirect when things go off track. Both matter, but they're very different mechanisms.

Lalam: Exactly. And the cultural implication is significant. As multi-agent AI systems become more common — in customer service, in healthcare, in education — we need frameworks for keeping them aligned with human goals. This paper provides one of the first empirical demonstrations that such frameworks can work.

Meng: And I'll give credit where it's due. Even with the simulation limitations, the engineering is solid. The fact that they ran sixty thousand simulations, tested across eight friction models and six personas, and still got a clean ninety-seven percent variance attribution — that's rigorous work.

Lu: The theoretical framing is also important. Calling it "adversarial-adjacent" — neither fully cooperative nor zero-sum — gives us a vocabulary for a whole class of problems we haven't been able to talk about precisely. That alone is a contribution.

Tom: So let's put a bow on this. The paper demonstrates that a control-theoretic governance layer can substitute for the missing goal function in multi-agent LLM systems. Two agents with opposed objectives can be steered toward a cooperative outcome without retraining either one.

Jane: And the key finding is that the governance policy, not the environment, determines where conversations end up. That's a hopeful message for anyone building these systems, because you can't control your users, but you can control how you respond to them.

Tom: The limitations are real — it's all simulation, and real humans are messier than any LLM persona. But the path forward is clear: live A/B testing, better reward signals, longer planning horizons, and testing across more domains.

Lalam: And if those next steps pan out, this framework could become a standard tool for anyone deploying multi-agent AI in the real world. That's a meaningful contribution to how AI integrates into society.

Tom: Well said, Lalam. That's a wrap on "Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes." Great paper, great discussion. Jane, what's coming up next?

Jane: Next up, we've got a paper on attention fine-tuning that this one actually references — it's about deriving reward signals from the model's own internal dynamics. Seems like a natural follow-up.

Tom: Perfect. We'll see you all then. Thanks for listening, everyone.

Alexander Liss, Nicholas Desmond, Santiago Gil Gallego

Georgia Institute of Technology · Huge Inc.

cs.AI

Submitted: 2026-04-25

Updated: 2026-08-13

Comments: 13 pages, 3 figures, 3 tables. Submitted to AI Engineer World's Fair 2026

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 46/100

Key concepts

Multi-LLM Agent Systems
This refers to scenarios where two or more AI agents operate simultaneously with distinct personas and goals. The paper addresses the challenge of these agents drifting or failing to reach a shared outcome because they lack a common mathematical target.
Experience Orchestrator (EO)
The EO is the governance layer designed to sit above interacting AI agents. It acts like a traffic controller, observing the conversation and applying corrective force to steer it toward a specific goal, rather than retraining the underlying AI models.
PID Controller
A feedback mechanism within the EO that monitors a visitor's resistance level. If resistance is high or stuck, this component applies corrective pressure or changes the strategy to keep the conversation moving toward success.
Contextual Bandit
This component acts as the decision-maker in the system. It selects one of four specific content strategies—such as direct pitch or trust-building guidance—based on all the information gathered about the visitor so far.

Terminology

Summary

Summary

This paper investigates whether a control theory-informed governance layer can substitute for the missing goal function in multi-agent LLM systems, steering two LLM agents with structurally opposed objectives toward a jointly optimal outcome. The authors state: "Classical RL agents optimize an explicit reward function; the goal is mathematically encoded in every gradient update. LLM agents have no equivalent. Their behavior is shaped by a prompt, which specifies intent in natural language but provides no formal optimization target. When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse [4]: neither agent has a mechanism to recognize the joint trajectory is suboptimal, and neither has a gradient signal to correct it [10], [11]. They frame the problem using Van Gelder's Dynamical Hypothesis and Kelso's coordination dynamics, noting that Multi-agent LLM systems exhibit the same behavior. Without external governance, the system drifts toward attractor states that are locally coherent but globally incoherent."

The proposed framework, the Experience Orchestrator (EO), governs the joint trajectory of two LLM agents in a simulated financial services environment. The environment models "a financial services website, where we model a human visitor browsing for retirement planning information and encountering an LLM-powered chatbot during their visit. The site agent seeks to guide the visitor toward a high-value action, scheduling a consultation with a financial advisor, while the visitor maintains realistic skepticism based on their assigned persona. The authors introduce the term adversarial-adjacent to describe this configuration: the two agents have structurally opposed turn-level objectives (conversion vs. resistance), but the system is designed around the hypothesis that governance can steer the joint trajectory toward a cooperative terminal state. By adversarial-adjacent we mean neither fully cooperative nor zero-sum: the agents do not share a reward function, but they are not competing for a fixed resource either."

EO operates through three governance mechanisms: "a PID controller that enforces behavioral consistency in real time; a POMDP belief tracker that maintains a probabilistic model of visitor intent; and a contextual bandit that selects the optimal content arm at each decision point. The state space consists of session-level features combined with a one-hot encoded page trajectory vector across ten pages, with the decision boundary at page p10 (Find an Advisor). The action space comprises four content arms: Contact (Speak with an Advisor), Guidance (Guidance Matters), Math (Time is on Your Side), and Questions (7 Questions to Ask"). The reward structure distinguishes between a binary terminal CB reward (Rterminal = 1 if the visitor selects the advisor contact arm with genuine resistance decline) and a composite shaped reward for per-turn feedback: Rt = wρ (−∆ρt) + wH (−∆Ht) + weff · 1t + wconv Rterminal with weights wρ = 0.3, wH = 0.2, weff = 0.1, wconv = 0.4.

The simulation was calibrated using SEM Rush web analytics data and included six domain-calibrated personas: digital native, fee hawk, legacy loyalist, stranded saver, dashboard exile, and grieving proxy. The experiment ran a full factorial design of 60,425 simulations spanning 8 friction models, and 6 persona archetypes. The winning variant, 'V4 SemRush', was compared against a Control Baseline defined as a naive LLM guiding the site agent that converses with the user and decides an action to take merely based on LLM reasoning and a system prompt, without any CB arm selection, PID control, or belief tracking.

The primary results are reported across three hypotheses. For lift: "V4 SemRush achieves a high-intent advisor contact rate of 78.1% versus 46.1% for Control Random (Naive LLM), a +32.0 point lift (Fisher's exact, p < 0.001). The governed system reaches 90% of the clairvoyant Oracle ceiling. For policy dominance: Two-way ANOVA across 60,425 simulations attributes 97% of between-factor outcome variance to CB variant selection and only 3% to friction model choice. The governing policy overwhelmingly determines where trajectories end up, regardless of environmental starting conditions. For trajectory quality: Under governance, arm selection aligns to persona reward priors and resistance declines monotonically. Without governance, arm mismatch escalates resistance until the dead-end detection mechanism terminates the session."

The per-persona analysis reveals two distinct regimes. In the persuasion-required regime, the Control LLM is essentially inert: digital native and fee hawk register near-zero baseline contact rates (0.6% and 1.0% respectively). The CB governance layer changes everything, delivering lifts of +68.7pp and +62.8pp. In the near-alignment regime, the naive LLM's empathetic defaults are already sufficient, and governance adds only marginal lift (+0.1pp to +3.9pp). The grieving proxy persona is "the critical exception. Control LLM performs at 90.1% as the unguided LLM naturally defaults to human contact for bereaved visitors; V4 SemRush drops to 65.6% (−24.5pp). The CB's structured arm-selection actively degrades the interaction by imposing a persuasion framework where empathy alone would suffice."

The authors identify three training failure modes: "The Sycophantic Collapse occurs when the visitor agent converges toward agreement without genuine engagement... The Stagnant Loop occurs when resistance neither rises nor falls... The Arm Fixation occurs when the CB concentrates on a single arm regardless of persona state. They note that All three are locally stable but globally suboptimal fixed points; the PID governance layer detects and corrects each."

The paper's central conclusion is that a control-theoretic governance layer can substitute for the missing goal function in a multi-agent LLM system, steering two agents with structurally differing objectives toward a jointly optimal outcome. The authors state: The governance layer has substituted for the missing goal function... effective multi-agent collaboration does not require re-architecting the agents or retraining them jointly. It requires building a better governor.

Key limitations are acknowledged: "All findings are conditional on LLM simulation... The visitor agent is another language model, not a human. Real visitors are far more unpredictable... The PID controller was calibrated against an LLM that reliably self-reports resistance scores on a structured scale; a human visitor provides no such signal. Future work includes Live A/B validation," HACA-Based Internal Reward Signal (augmenting the PID heuristic with a self-supervised reward from decoder cross-attention dynamics), Domain generalization to healthcare, enterprise software sales, HR recruiting, and technical support, and POMCP Lookahead Planning to simulate 3–5 turns ahead.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and what the improved AI system can do:

  • Implementation: Wrap two or more LLM agents (e.g., a persuader and a resister) with a PID controller that observes a quantifiable state (e.g., user resistance, engagement score, task progress) and dynamically constrains the agents' structured output (e.g., JSON schema bounds) each turn.

  • What it does: Prevents the common failure mode where LLM agents drift toward sycophantic agreement or get stuck in polite, unproductive loops. The system now actively steers the joint trajectory toward a target outcome (e.g., conversion, task completion) without retraining the models.

  • Implementation: Add a lightweight contextual bandit (e.g., LightGBM with UCB) that, given session context (page trajectory, device, channel, persona signals), selects one of several content arms (e.g., Contact, Guidance, Math, Questions) before each agent turn.

  • What it does: Replaces naive LLM reasoning for what to say next with data-driven arm selection. In the paper, this alone accounted for 97% of outcome variance, delivering a +32-point lift in high-intent conversion versus a prompt-only LLM. The system now adapts its persuasion/guidance strategy per user context rather than guessing.

  • Implementation: Maintain a Dirichlet distribution over user intent states (e.g., browse, compare, purchase, support). Update it each turn using keyword signals from the user's message. Feed the belief entropy and distribution into the PID gain scheduler and the bandit context.

  • What it does: The system now explicitly models uncertainty about user intent. When uncertainty is high, it loosens control (avoids over-aggressive persuasion); when intent becomes clear, it tightens control and selects the most effective arm. This prevents misaligned actions (e.g., pushing a sale to a user who just wants information).

  • Implementation: Monitor resistance/engagement trends using EMA and SMA. If resistance is unchanged for 3+ turns (stagnation) or if a mismatched arm escalates resistance, automatically switch arms or gracefully terminate the conversation.

  • What it does: Prevents wasted interactions and user frustration. The system now recognizes when it is trapped in a locally stable but globally suboptimal state and self-corrects, rather than continuing to repeat the same ineffective message.

  1. Run a two-agent conversation (e.g., sales bot + skeptical user) that reliably reaches a high-value terminal action (e.g., booking a consultation, completing a form) at a rate 32 percentage points higher than a naive prompt-only LLM, while filtering out sycophantic false positives.

  2. Adapt its persuasion strategy in real time — if a user resists a direct pitch, it pivots to a trust-building or evidence-based arm; if the user is analytical, it leads with quantitative data; if the user is anxious, it uses low-commitment framing.

  3. Explicitly track and react to user intent uncertainty — it does not over-persuade when unsure, and it tightens its approach once intent becomes clear.

  4. Self-detect and escape conversational dead-ends — if the user is stuck in a loop or resistance is not moving, the system changes strategy or ends the conversation gracefully, avoiding wasted time and negative user experience.

  5. Generalize across user personas — the governance layer works for different archetypes (skeptical, anxious, loyal, overwhelmed) without retraining the underlying LLM, by adjusting arm selection and PID gains per persona.

Bottom line: This system is not just a smarter chatbot; it is a governed multi-agent system that uses control theory to substitute for a missing reward function, delivering measurable, variance-dominant improvements in collaborative outcomes.

Sources

Related papers