When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

arXiv:2608.07538 · cs.AI, cs.GT, econ.GN, q-fin.EC · Submitted 2026-07-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains".

Jane: The paper was written by Chen Liang and Fasheng Xu from University of Connecticut.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, listeners, welcome back. Today we're cracking open a paper that's been making waves in the operations world, and it's got a title that just rolls off the tongue: "When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains." Jane, I have to say, just reading that title gets me excited.

Jane: Oh, me too, Tom. And honestly, the title tells you exactly what's at stake here. We're not talking about a chatbot writing a poem. We're talking about AI agents sitting across the table from each other, haggling over contracts. The paper is asking a really fundamental question: if you hand the negotiation over to an AI, are you getting a good deal?

Tom: Right, and it's not just any negotiation. It's a supply chain problem where one side knows something the other doesn't. The buyer knows their own demand, the seller doesn't. That's the "private information" part. It's like playing poker where one player can see their own cards but the other can only guess.

Jane: And that's what makes it so clever. They're using a classic economic model, the one from Feng and his colleagues, as a yardstick. They set up thousands of these AI-versus-AI negotiations and then check the results against what a perfectly rational, mathematically perfect negotiator would do. It's a benchmark test.

Tom: A benchmark test with real teeth. I mean, they ran nearly ten thousand negotiations. That's not a small sample. And they're using models from OpenAI, Google, and Alibaba. So you're getting a real cross-section of the industry's best minds, so to speak.

Jane: Exactly. And the title hints at the big finding, which we'll get into, but the core idea is that these agents are actually pretty good at creating value, but they're not always great at dividing it up fairly or predictably. And who you pick to be your agent matters a lot.

Tom: It's like hiring a lawyer. A great lawyer might win you the case, but a different great lawyer might settle for a different amount. The paper is saying that the "personality" of the AI, which comes from the company that built it, is a huge factor in how the money gets split.

Jane: So for a company deciding to automate their procurement, this isn't just a tech choice. It's a strategic business decision. You're essentially choosing your bargaining style when you pick your vendor. That's the headline.

Tom: And that's the hook. We're going to spend the rest of the show digging into how they figured that out, what the numbers actually say, and what it means for the future of commerce. Stick around.

Summary: Tom: So, Jane, we've set the stage. Let's get into the meat of this paper, "When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains." What did they actually find?

Jane: Well, the headline is that these agents are shockingly good at closing a deal. They reached an agreement in almost ninety-nine percent of the negotiations. And when they did agree, they captured over ninety-five percent of the theoretical maximum surplus. That's the total pie of value available. So they're not leaving much money on the table in terms of the final contract.

Tom: But here's the rub. They take their sweet time doing it. The theoretically perfect negotiator settles in about one point two five rounds. These agents average almost three rounds. And in a negotiation, time is money. The paper shows that this delay erodes between twenty-one percent and thirty-four percent of that first-best surplus, depending on how you value patience.

Lu: That's the key insight from my perspective. The agents are good at finding the efficient outcome, but they're inefficient in the process of getting there. They're not doing the one-shot, perfectly calculated move. They're engaging in a kind of iterative search, proposing, getting rejected, and counter-proposing.

Jane: Exactly, Lu. And that's where the capability of the model comes in. The top-tier models are more efficient in the final outcome, but they actually take more rounds than the weaker models. It's like a chess grandmaster thinking for twenty minutes versus a club player moving in five seconds. The grandmaster finds the better move, but it takes longer.

Tom: And that delay is a real cost. The paper calls it "discounted efficiency." It's not just about the final contract; it's about how long it took to get there. And that's a crucial lesson for anyone deploying these agents.

Meng: From an engineering standpoint, that's a huge red flag. If I'm building a system to run thousands of these negotiations, a two point nine eight-round average versus a one point two five-round benchmark means I need to budget for a lot more API calls and a lot more latency. The cost of that delay is baked into the system's operational expenses.

Tom: So the value creation is there, but the process is costly. And that's just the first finding. The second one is where it gets really wild. It's not just about how capable the model is; it's about who made it. The paper found that the provider of the model is a better predictor of who wins the negotiation than the model's capability.

Jane: Right. So if you're using Alibaba's Qwen models, you're going to get a much better deal as the buyer than if you're using OpenAI's models. The paper shows a huge swing in surplus share just based on the vendor. It's like the AI has a built-in bias for one side of the table.

Lu: And that's a profound finding. It suggests that the training data and the alignment process, which are different for each company, bake in a certain "bargaining personality." It's not something you can prompt your way out of easily. It's a property of the model itself.

Tom: So, the takeaway is that choosing your AI vendor is choosing your negotiation strategy. We'll get into the cross-provider matchups next, because that's where this really gets fascinating.

Improvements: Tom: So we've established that these agents are good at making the pie, but the provider decides how it's sliced. Now, what does the paper suggest we actually do about it? What are the "improvements" or the levers we can pull?

Jane: Right. The paper is not just a doom-and-gloom report. It's a guide for how to deploy these agents responsibly. And the first big lever is the prompt itself. You can change the agent's behavior just by changing the instructions you give it.

Tom: And we're not talking about small tweaks. The paper shows that if you restrict the agents to only sending numbers, no natural language, the surplus division shifts. For some models, it shifts a lot. So the ability to chat is doing real strategic work. It's not just decoration.

Lu: That's a fascinating point. It means the verbal channel is a tool for persuasion and information hiding. When you take it away, the agents have to rely purely on the structure of the offer. And some models are better at that than others. The paper shows that removing the verbal channel changes the buyer's share by a significant margin for certain providers.

Meng: So as an engineer, that tells me the interface matters as much as the model. If I'm building a procurement system, I have a choice: do I let my agent talk freely, or do I constrain it to a structured format? The paper is saying that choice has a direct impact on the bottom line.

Jane: Exactly. And the second lever is even more interesting. It's about patience. In classical economics, patience is a fixed trait. But here, the principal can choose the agent's "strategic patience" by putting a number in the prompt. You can tell your agent to be patient or to be in a hurry, regardless of your own real-world time constraints.

Tom: And that's a game-changer. The paper calls this the separation of "economic patience" from "strategic patience." Your real cost of delay is one thing, but the patience you program into your agent is another. And the paper shows that choosing the right strategic patience can have a huge impact on your payoff.

Lu: It's a beautiful insight. You're no longer bound by your own psychology. You can deploy an agent that is more patient than you are, or less, depending on the situation. The paper's data shows that for some models, being less patient as a buyer actually gets you a better deal. It's counterintuitive, but it works.

Jane: And the third lever is about guardrails. The paper found that the weakest models will accept deals that lose money. That's a huge operational risk. So the suggestion is that if you're using a baseline model, you absolutely need an automated system to check the profitability of every deal before it's accepted.

Tom: So it's not just about picking the smartest model. It's about configuring the system correctly: the prompt, the patience, and the safety checks. The paper is giving us a playbook for making this work.

Meng: And that playbook is exactly what we need. It moves this from a cool experiment to something we can actually build and deploy with confidence.

First Page: Tom: So we've got the playbook. But let's go back to the very beginning of the paper, the first page, because it sets up the stakes so well. It opens with the real-world examples that make this so urgent.

Jane: It does. It talks about Walmart using LLM agents for thousands of supplier contracts. It mentions Alibaba's tools for sourcing agents. And it even brings up an experiment where AI agents were trading real personal items. This isn't a theoretical exercise. This is happening right now.

Tom: And the paper makes a really sharp point about that. When both sides of a transaction are automated, you get LLM-to-LLM negotiation happening at machine speed. There was even a competition mentioned with over one hundred eighty thousand AI-to-AI negotiations. That's a scale that's impossible for humans.

Lu: And that scale is exactly why the paper's findings are so important. If you have a systematic bias in one provider's models, it's not just one bad deal. It's thousands of bad deals, all happening automatically. The errors and biases scale up with the automation.

Jane: That's the real danger. A human negotiator might have a bad day, but an AI with a built-in bias will have a bad day every single time, for every single contract. The paper is essentially saying we need to audit these agents before we let them loose.

Meng: And the first page also frames the core question perfectly. It's not about whether these agents can talk like humans. It's about whether they can advance their principal's economic interests. Can they actually get you a good deal? That's the only question that matters for a business.

Tom: And that's what the rest of the paper answers. It gives you the tools to measure that. It provides a benchmark, the Perfect Bayesian Equilibrium, to compare against. It's a way to say, "My agent got ninety-five percent of the surplus, but the theoretical max was one hundred percent. Is that good enough?"

Jane: And the answer, as we've seen, is that it depends. It depends on the model, the provider, and how you configure the prompt. The first page sets up this whole framework for thinking about it, which is why it's so well-written.

Lu: It really is. It takes a complex problem and frames it in a way that's immediately actionable. It's not just an academic exercise. It's a call to action for anyone building or using these systems.

Tom: So we've got the problem, the framework, and the findings. Let's wrap this up and see what it all means for the future.

Conclusion: Tom: Well, Jane, we've been through the whole journey with "When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains." Let's try to tie it all together.

Jane: Let's do it. The paper gives us three big lessons. First, capability is about creating value. The best models are great at finding the most efficient contract, but they're slow, and that delay costs money. Second, provider identity is about dividing value. Who you choose as your AI vendor has a massive, predictable impact on which side of the table gets the better deal.

Tom: And the third lesson is that you, the person deploying the agent, have more control than you think. You can change the prompt, you can change the agent's patience, and you can add guardrails. The paper shows that these configuration choices are just as important as the model choice itself.

Lu: And I think the most profound implication is that we can no longer think of AI agents as neutral tools. They have distinct "bargaining personalities" that are shaped by their creators. This paper gives us the methodology to measure and understand those personalities.

Meng: From my side, it means we can build systems that are not just powerful, but also predictable and safe. We can test our agents against this benchmark, we can tune their prompts, and we can put in the verification layers to catch the mistakes. It turns a gamble into an engineering problem.

Jane: And that's the real takeaway. This paper isn't just a warning. It's a guide. It's a way to make sure that when you delegate your negotiation to an AI, you're getting the deal you wanted, not the deal the AI's training data decided you should get.

Tom: It's a fantastic piece of research, and it's going to be a reference point for anyone working in this space. So, with that, we're going to say goodbye to "When LLM Agents Negotiate." It's been a real eye-opener.

Jane: It really has. Thanks for joining us, everyone. We'll be back soon with another paper that's shaping the future of AI. Until then, take care.

Chen Liang, Fasheng Xu

University of Connecticut

cs.AI, cs.GT, econ.GN, q-fin.EC

Submitted: 2026-07-29

Updated: 2026-08-11

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

The gist: This paper studies whether general-purpose LLM agents can recover equilibrium-consistent bargaining behavior in a canonical supply chain contracting problem, and how their behavior varies with the

Key concepts

Private Information
This occurs when one party in a negotiation knows something the other does not. In supply chains, the buyer knows their own demand, but the seller does not. This asymmetry is central to how these AI agents negotiate and bargain over contracts.
Discounted Efficiency
This refers to the cost of time in a negotiation. While LLM agents are good at finding the optimal contract, they take significantly more rounds than a theoretically perfect negotiator. This delay erodes potential economic surplus.
AI Bargaining Personality
The provider of the LLM agent (e.g OpenAI vs Alibaba) has a built-in bias that influences the negotiation outcome, regardless of the model's capability. This suggests that choosing your AI vendor is essentially choosing your bargaining strategy in business deals.
Strategic Patience
This is a programmable trait where the user dictates how patient the AI agent should be, separate from real-world time constraints. The paper shows that setting this strategic patience can significantly impact the final payoff and deal structure.

Terminology

Summary

This paper studies whether general-purpose LLM agents can recover equilibrium-consistent bargaining behavior in a canonical supply chain contracting problem, and how their behavior varies with the vendor-selection and role-assignment decisions firms must make. The authors benchmark nine LLMs from three providers (OpenAI, Google, and Alibaba) against the Perfect Bayesian Equilibrium of the alternating-offer bargaining game of Feng et al. (2015), in which a buyer with private demand information negotiates a quantity–payment contract (q, T) with an uninformed seller. The experimental program comprises 9,840 LLM-to-LLM negotiations across three main-analysis blocks (symmetric verbal self-play, cross-family flagship pairings, and within-family capability-asymmetric pairings) and seven design and robustness extensions (structured-offer bargaining without verbal communication, no-discounting bargaining, strategic-patience analysis, reasoning-effort ablation, retail-price sensitivity, prior sensitivity, and parameter-size variation). Prompts disclose the primitives of the game (payoff formulas, demand distributions, discount factors, priors, and buyer type) but agents are not told which contract to offer, when to separate types, or how to update beliefs. The design therefore audits strategic execution in a stylized supply chain contracting setting rather than discovery of the setting from raw context.

The paper delivers three empirical findings tied to the vendor-selection and role-assignment problem.

Finding 1: Capability is the value-creation lever that governs both efficiency and reliability. LLM agents reach agreement in 98.9% of negotiations and capture 95.4% of first-best surplus in undiscounted terms, but they average 2.98 rounds against the Bayesian benchmark of 1.25. Under discounting—which we treat as a stress test on costly continuation rather than a literal time-preference estimate—this delay erodes 21–34% of first-best surplus, depending on patience. Of the efficiency variation the experimental design systematically moves, capability explains roughly 66% (Shapley decomposition): flagship models take more rounds than baselines (3.25 versus 2.75) but achieve materially higher efficiency (98.9% versus 91.0%), consistent with an iterative-search pattern in which capable agents invest in proposal–rejection cycles rather than settling at round-one heuristic offers. The same capability ordering governs operational reliability: baseline models accept individually irrational contracts (negative profit for one party) in 19.2% of cases, versus 0.6% for mid-tier and 0.0% for flagship models. This order-of-magnitude gap separates a verification-gated weak-model cluster from a lighter-monitoring strong-model cluster, making automated profit verification a necessary guardrail below the threshold.

Finding 2: Provider identity is the distributional lever, and surplus capture depends on the counterparty and role. Capability and provider identity load on different margins: capability governs how much surplus is created (Finding 1), while provider identity is the more reliable predictor of who captures it. Among the model versions evaluated under a common prompt protocol, self-play surplus division varies markedly across providers: Qwen models average 70% buyer share (within-family range 53–91% across capability tiers), Gemini 50% (44–54%), and OpenAI 40% (38–42%). The mean gap between Qwen and OpenAI (approximately 30 percentage points) is comparable to the largest within-provider spread (39 points across Qwen tiers), so the 53.5% pooled buyer share masks three qualitatively distinct bargaining profiles rather than a single capability-tier effect. Of the surplus-division variation the experimental design systematically moves, the announced patience parameters explain 90% (Shapley decomposition), while capability, buyer type, and first-proposer assignment together contribute the remaining 10%. Prompted, common-knowledge patience, however, is a configuration choice rather than a model attribute (Finding 3); among the model-side factors a firm cannot prompt away, provider identity is the strongest predictor of who captures surplus. Public demonstrations such as Anthropic’s Project Deal, in which upgrading one’s own agent reliably improved its terms, reinforce the intuition that a more capable agent is a better bargainer. Our results qualify it: capability reliably creates value (Finding 1), but its distributional advantage does not survive across providers. Even holding provider fixed, a cross-tier GPT pairing makes the point: switching the buyer from GPT-5.2 to the weaker GPT-5-mini against an unchanged GPT-5.2 seller lowers buyer share from 37.9% to 35.2%, so within a family the more capable model captures the larger share. Across providers, however, capability rank alone does not determine who captures surplus: a highly capable Qwen flagship is the weakest cross-family seller, so a model’s provider profile can override its capability. Cross-family direction effects amplify this pattern: when flagship models from different providers negotiate each other, the direction of the pairing produces 7–18 percentage-point swings in surplus, comparable to the within-family capability swing; Gemini captures 67% of surplus as a cross-family buyer and retains 48% as seller, while Qwen retains only 27% as a cross-family seller. For firms deploying different vendors on opposite sides of a procurement interaction, model choice is therefore a first-order distributional decision, not a commodity input.

Finding 3: The principal’s configuration choices are strategic levers. Because firms bargain through delegated agents rather than directly, the prompt serves as the agent’s mandate: three economically meaningful design choices systematically shape bargaining outcomes. Constraining the communication channel reshapes provider biases: removing natural language shifts OpenAI’s baseline and mid-tier models toward buyers (GPT-5-mini: 41.9% to 51.3%) while shifting Gemini’s mid-tier and flagship models toward sellers, so the verbal channel is doing strategic work, not decorating offer behavior. Removing the discounting framework from the prompt leaves the value-creation, distributional, and reliability rankings intact but changes negotiation tempo, so discount-factor language operates as a tempo cue rather than a substantive driver of the results. Most distinctively, delegation to LLM agents separates two objects that classical bargaining treats as a single discount-factor parameter: economic patience (the principal’s fixed real-time cost of delay) and strategic patience (the agent’s prompted discount factor, which the principal sets at deployment). Patience thus becomes a configurable design lever rather than an exogenous primitive, and choosing it well raises realized payoff by economically meaningful margins. Configuration is therefore a first-order design problem alongside model and vendor selection.

The paper makes three contributions to operations management research on AI-mediated contracting. Empirically, we document a dimension of LLM heterogeneity that standard capability benchmarks miss (provider-level bargaining profiles) and show it is as consequential as within-family capability differences once counterparties are heterogeneous, establishing vendor choice as a strategic operations decision rather than a technical one. Methodologically, we introduce a reproducible equilibrium-referenced audit framework: an executable implementation of the branch-specific PBE, validated cell-by-cell against the equilibrium’s analytical properties—a portable template for evaluating autonomous agents in game-theoretic operational settings. Because the audit spans the current capability frontier (nine models, three tiers, and roughly three generations), it separates structural regularities that travel across model vintages from model-specific results that recalibrate as providers update, so the framework remains informative as the underlying models improve. Practically, we identify three deployment-relevant evaluation dimensions (time-adjusted efficiency, provider-level distributional profile, and operational reliability), guiding capability thresholds, role assignment, and guardrail architecture.

Delegation changes both the stakes and the relevant benchmark. When firms authorize LLM agents to negotiate on their behalf, agents’ offers, concessions, and acceptance decisions directly determine realized contract terms and surplus. This is particularly consequential in procurement, where automation can extend bargaining to supplier relationships that firms cannot economically negotiate one by one, allowing gains and systematic errors alike to scale across transactions. The central question is therefore not whether LLM agents bargain like humans, but whether they advance their principals’ economic interests. We evaluate this competence against the Perfect Bayesian Equilibrium characterized by Feng et al. (2015). The equilibrium provides reference values for agreement timing, contract form, surplus division, screening, and feasibility. Theory thus serves as measurement infrastructure, allowing us to identify when an agent creates value, transfers value to its counterparty, or destroys value through delay or individually irrational acceptance. Human-subject experiments address complementary questions about human–agent interaction and the transmission of principal heterogeneity into delegated outcomes; for example, Imas et al. (2025) show that delegated outcomes retain substantial variation across principals. Our design instead fixes the payoff objective and evaluates the agent against the corresponding equilibrium, providing a direct measure of agent-level bargaining competence.

The paper concludes with a three-dimensional deployment framework comprising three deployment rules that follow directly from the findings. Time-adjusted efficiency shifts the locus of control from outcome to process: because deal-level outcomes carry meaningful stochasticity even under flagship models, organizations should bound delay ex ante through round caps, escalation rules for stalled negotiations, and final-offer procedures for high-urgency contracts. Distributional profile makes vendor choice first-order whenever counterparties run different providers: cross-family direction effects (7–18 pp) are comparable to within-family tier effects (∼17 pp), so firms controlling both sides of a transaction should weight provider identity at least as heavily as capability rank when assigning roles, and, because self-play profiles do not transfer directly to heterogeneous matchups, audit the specific cross-vendor pairing in simulation before delegation. Operational reliability separates into two regimes: the 0.0–0.6% irrationality rate at flagship and mid-tier permits lighter monitoring, while the 19.2% rate at baseline makes dual-party automated profit verification non-negotiable. Because arithmetic errors, failed constraint checks, and instruction non-compliance are operationally equivalent (each yields an economically unsafe contract), the verification layer should treat them uniformly.

The paper distinguishes structural regularities (high undiscounted efficiency with delayed agreement, provider-level heterogeneity in surplus division, and concentration of unsafe agreements among weaker models) from model-specific results such as specific surplus shares, irrationality rates, and provider rankings. The former are likely to travel across nearby settings; the latter are calibrated to the model versions tested here and require periodic recalibration as providers update their systems. This division rests on evidence, not assertion: across the three capability tiers and roughly three model generations we test, the structural regularities recur at every point along the frontier, including the reliability gradient traced continuously by the R7 capacity ladder, while only the model-specific results shift with capability. Extending the structural claims to models outside this tested range is a prediction the design supports but does not itself verify.

Three limitations bound the interpretation of these results. First, prompts are role-specific, so the reported surplus shares reflect behavior under a common prompting regime and are not prompt-free estimates of intrinsic bargaining bias; the qualitative pairing order is broadly stable across our prompt variations, with specific exceptions identified in Section 6.4. Second, our design discloses each party’s patience parameters as common knowledge to both agents, matching the informational structure of the Bayesian benchmark; whether the qualitative patterns persist when a principal’s patience or urgency is undisclosed to the counterparty, a common feature of real negotiations, is untested. Third, several interpretive moves remain process-descriptive rather than mechanistically identified: the strategic-behavior evidence of Section 4.3 documents heuristic substitution and a behavior–theory gap in equilibrium comparative statics without adjudicating among competing cognitive explanations, and the hazard-based reading of δ in Section 3.3 is an analogy rather than a tested mechanism.

These limitations point to a concrete research agenda. The hazard-based reading of δ is testable by exogenously varying per-round reliability (through compute budgets, injected parsing noise, or forced-termination protocols), and the related conjecture that effective δ declines as conversation history accumulates can be tested by measuring error rates against round number across context-management regimes. If these mappings hold, engineering choices organizations already make carry bargaining consequences (round budgets as value-conditional commitment devices, profit-verification guardrails as credible commitments, and an agent-versus-human boundary drawn along per-interaction reliability rather than general capability). Field studies, in turn, would assess which structural regularities survive once agents are embedded in real approval and verification workflows. As LLM agents move from pilot to production in procurement, the operational question shifts from whether a delegated agent can close a deal to whether the deal it closes is one the principal would have authorized. The audit developed here answers that question before delegation, not after—replacing the intuition that a more capable agent is simply a better bargainer with a discipline that treats capability, provider, and configuration as three separate levers a principal must each get right.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

Improvement: Integrate a validated Perfect Bayesian Equilibrium (PBE) benchmark into LLM negotiation systems as a real-time reference layer. The system computes branch-specific equilibrium contracts, agreement timing, and surplus division for the current bargaining context.

What the improved system can do:

  • Detect when an agent's proposed contract deviates from the equilibrium benchmark (e.g., quantity distortion, suboptimal payment) and flag it for correction

  • Estimate the efficiency loss from delayed agreement in real time, alerting the principal when rounds exceed the equilibrium prediction (1.25 rounds in the tested setting)

  • Automatically verify that accepted contracts yield non-negative expected profit for the principal, preventing the 19.2% irrational-agreement rate seen in baseline models

Sources

Related papers