When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains
summary
The gist
This paper studies whether general-purpose LLM agents can recover equilibrium-consistent bargaining behavior in a canonical supply chain contracting problem, and how their behavior varies with the
In short
The episode discusses a study on LLM agents negotiating complex supply chain contracts. While these AI agents are highly effective at closing deals and capturing value, the hosts conclude that the negotiation process is slow. Crucial findings include how much the AI provider's identity affects deal division and how users can control outcomes via prompt configuration.
Key concepts
- Private Information
- This occurs when one party in a negotiation knows something the other does not. In supply chains, the buyer knows their own demand, but the seller does not. This asymmetry is central to how these AI agents negotiate and bargain over contracts.
- Discounted Efficiency
- This refers to the cost of time in a negotiation. While LLM agents are good at finding the optimal contract, they take significantly more rounds than a theoretically perfect negotiator. This delay erodes potential economic surplus.
- AI Bargaining Personality
- The provider of the LLM agent (e.g OpenAI vs Alibaba) has a built-in bias that influences the negotiation outcome, regardless of the model's capability. This suggests that choosing your AI vendor is essentially choosing your bargaining strategy in business deals.
- Strategic Patience
- This is a programmable trait where the user dictates how patient the AI agent should be, separate from real-world time constraints. The paper shows that setting this strategic patience can significantly impact the final payoff and deal structure.
Terminology used across episodes
This episode discusses
- When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains · Paper Radio
- Assured autonomy: How operations research powers and orchestrates generative AI systems
- An Economy of AI Agents
- Large Language Newsvendor: Decision Biases and Cognitive Mechanisms
- Generative AI and Organizational Structure in the Knowledge Economy
The paper
When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains · Read on arXiv
Chen Liang, Fasheng Xu
University of Connecticut
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains".
Jane: The paper was written by Chen Liang and Fasheng Xu from University of Connecticut.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, listeners, welcome back. Today we're cracking open a paper that's been making waves in the operations world, and it's got a title that just rolls off the tongue: "When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains." Jane, I have to say, just reading that title gets me excited.
Jane: Oh, me too, Tom. And honestly, the title tells you exactly what's at stake here. We're not talking about a chatbot writing a poem. We're talking about AI agents sitting across the table from each other, haggling over contracts. The paper is asking a really fundamental question: if you hand the negotiation over to an AI, are you getting a good deal?
Tom: Right, and it's not just any negotiation. It's a supply chain problem where one side knows something the other doesn't. The buyer knows their own demand, the seller doesn't. That's the "private information" part. It's like playing poker where one player can see their own cards but the other can only guess.
Jane: And that's what makes it so clever. They're using a classic economic model, the one from Feng and his colleagues, as a yardstick. They set up thousands of these AI-versus-AI negotiations and then check the results against what a perfectly rational, mathematically perfect negotiator would do. It's a benchmark test.
Tom: A benchmark test with real teeth. I mean, they ran nearly ten thousand negotiations. That's not a small sample. And they're using models from OpenAI, Google, and Alibaba. So you're getting a real cross-section of the industry's best minds, so to speak.
Jane: Exactly. And the title hints at the big finding, which we'll get into, but the core idea is that these agents are actually pretty good at creating value, but they're not always great at dividing it up fairly or predictably. And who you pick to be your agent matters a lot.
Tom: It's like hiring a lawyer. A great lawyer might win you the case, but a different great lawyer might settle for a different amount. The paper is saying that the "personality" of the AI, which comes from the company that built it, is a huge factor in how the money gets split.
Jane: So for a company deciding to automate their procurement, this isn't just a tech choice. It's a strategic business decision. You're essentially choosing your bargaining style when you pick your vendor. That's the headline.
Tom: And that's the hook. We're going to spend the rest of the show digging into how they figured that out, what the numbers actually say, and what it means for the future of commerce. Stick around.
Summary: Tom: So, Jane, we've set the stage. Let's get into the meat of this paper, "When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains." What did they actually find?
Jane: Well, the headline is that these agents are shockingly good at closing a deal. They reached an agreement in almost ninety-nine percent of the negotiations. And when they did agree, they captured over ninety-five percent of the theoretical maximum surplus. That's the total pie of value available. So they're not leaving much money on the table in terms of the final contract.
Tom: But here's the rub. They take their sweet time doing it. The theoretically perfect negotiator settles in about one point two five rounds. These agents average almost three rounds. And in a negotiation, time is money. The paper shows that this delay erodes between twenty-one percent and thirty-four percent of that first-best surplus, depending on how you value patience.
Lu: That's the key insight from my perspective. The agents are good at finding the efficient outcome, but they're inefficient in the process of getting there. They're not doing the one-shot, perfectly calculated move. They're engaging in a kind of iterative search, proposing, getting rejected, and counter-proposing.
Jane: Exactly, Lu. And that's where the capability of the model comes in. The top-tier models are more efficient in the final outcome, but they actually take more rounds than the weaker models. It's like a chess grandmaster thinking for twenty minutes versus a club player moving in five seconds. The grandmaster finds the better move, but it takes longer.
Tom: And that delay is a real cost. The paper calls it "discounted efficiency." It's not just about the final contract; it's about how long it took to get there. And that's a crucial lesson for anyone deploying these agents.
Meng: From an engineering standpoint, that's a huge red flag. If I'm building a system to run thousands of these negotiations, a two point nine eight-round average versus a one point two five-round benchmark means I need to budget for a lot more API calls and a lot more latency. The cost of that delay is baked into the system's operational expenses.
Tom: So the value creation is there, but the process is costly. And that's just the first finding. The second one is where it gets really wild. It's not just about how capable the model is; it's about who made it. The paper found that the provider of the model is a better predictor of who wins the negotiation than the model's capability.
Jane: Right. So if you're using Alibaba's Qwen models, you're going to get a much better deal as the buyer than if you're using OpenAI's models. The paper shows a huge swing in surplus share just based on the vendor. It's like the AI has a built-in bias for one side of the table.
Lu: And that's a profound finding. It suggests that the training data and the alignment process, which are different for each company, bake in a certain "bargaining personality." It's not something you can prompt your way out of easily. It's a property of the model itself.
Tom: So, the takeaway is that choosing your AI vendor is choosing your negotiation strategy. We'll get into the cross-provider matchups next, because that's where this really gets fascinating.
Improvements: Tom: So we've established that these agents are good at making the pie, but the provider decides how it's sliced. Now, what does the paper suggest we actually do about it? What are the "improvements" or the levers we can pull?
Jane: Right. The paper is not just a doom-and-gloom report. It's a guide for how to deploy these agents responsibly. And the first big lever is the prompt itself. You can change the agent's behavior just by changing the instructions you give it.
Tom: And we're not talking about small tweaks. The paper shows that if you restrict the agents to only sending numbers, no natural language, the surplus division shifts. For some models, it shifts a lot. So the ability to chat is doing real strategic work. It's not just decoration.
Lu: That's a fascinating point. It means the verbal channel is a tool for persuasion and information hiding. When you take it away, the agents have to rely purely on the structure of the offer. And some models are better at that than others. The paper shows that removing the verbal channel changes the buyer's share by a significant margin for certain providers.
Meng: So as an engineer, that tells me the interface matters as much as the model. If I'm building a procurement system, I have a choice: do I let my agent talk freely, or do I constrain it to a structured format? The paper is saying that choice has a direct impact on the bottom line.
Jane: Exactly. And the second lever is even more interesting. It's about patience. In classical economics, patience is a fixed trait. But here, the principal can choose the agent's "strategic patience" by putting a number in the prompt. You can tell your agent to be patient or to be in a hurry, regardless of your own real-world time constraints.
Tom: And that's a game-changer. The paper calls this the separation of "economic patience" from "strategic patience." Your real cost of delay is one thing, but the patience you program into your agent is another. And the paper shows that choosing the right strategic patience can have a huge impact on your payoff.
Lu: It's a beautiful insight. You're no longer bound by your own psychology. You can deploy an agent that is more patient than you are, or less, depending on the situation. The paper's data shows that for some models, being less patient as a buyer actually gets you a better deal. It's counterintuitive, but it works.
Jane: And the third lever is about guardrails. The paper found that the weakest models will accept deals that lose money. That's a huge operational risk. So the suggestion is that if you're using a baseline model, you absolutely need an automated system to check the profitability of every deal before it's accepted.
Tom: So it's not just about picking the smartest model. It's about configuring the system correctly: the prompt, the patience, and the safety checks. The paper is giving us a playbook for making this work.
Meng: And that playbook is exactly what we need. It moves this from a cool experiment to something we can actually build and deploy with confidence.
First Page: Tom: So we've got the playbook. But let's go back to the very beginning of the paper, the first page, because it sets up the stakes so well. It opens with the real-world examples that make this so urgent.
Jane: It does. It talks about Walmart using LLM agents for thousands of supplier contracts. It mentions Alibaba's tools for sourcing agents. And it even brings up an experiment where AI agents were trading real personal items. This isn't a theoretical exercise. This is happening right now.
Tom: And the paper makes a really sharp point about that. When both sides of a transaction are automated, you get LLM-to-LLM negotiation happening at machine speed. There was even a competition mentioned with over one hundred eighty thousand AI-to-AI negotiations. That's a scale that's impossible for humans.
Lu: And that scale is exactly why the paper's findings are so important. If you have a systematic bias in one provider's models, it's not just one bad deal. It's thousands of bad deals, all happening automatically. The errors and biases scale up with the automation.
Jane: That's the real danger. A human negotiator might have a bad day, but an AI with a built-in bias will have a bad day every single time, for every single contract. The paper is essentially saying we need to audit these agents before we let them loose.
Meng: And the first page also frames the core question perfectly. It's not about whether these agents can talk like humans. It's about whether they can advance their principal's economic interests. Can they actually get you a good deal? That's the only question that matters for a business.
Tom: And that's what the rest of the paper answers. It gives you the tools to measure that. It provides a benchmark, the Perfect Bayesian Equilibrium, to compare against. It's a way to say, "My agent got ninety-five percent of the surplus, but the theoretical max was one hundred percent. Is that good enough?"
Jane: And the answer, as we've seen, is that it depends. It depends on the model, the provider, and how you configure the prompt. The first page sets up this whole framework for thinking about it, which is why it's so well-written.
Lu: It really is. It takes a complex problem and frames it in a way that's immediately actionable. It's not just an academic exercise. It's a call to action for anyone building or using these systems.
Tom: So we've got the problem, the framework, and the findings. Let's wrap this up and see what it all means for the future.
Conclusion: Tom: Well, Jane, we've been through the whole journey with "When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains." Let's try to tie it all together.
Jane: Let's do it. The paper gives us three big lessons. First, capability is about creating value. The best models are great at finding the most efficient contract, but they're slow, and that delay costs money. Second, provider identity is about dividing value. Who you choose as your AI vendor has a massive, predictable impact on which side of the table gets the better deal.
Tom: And the third lesson is that you, the person deploying the agent, have more control than you think. You can change the prompt, you can change the agent's patience, and you can add guardrails. The paper shows that these configuration choices are just as important as the model choice itself.
Lu: And I think the most profound implication is that we can no longer think of AI agents as neutral tools. They have distinct "bargaining personalities" that are shaped by their creators. This paper gives us the methodology to measure and understand those personalities.
Meng: From my side, it means we can build systems that are not just powerful, but also predictable and safe. We can test our agents against this benchmark, we can tune their prompts, and we can put in the verification layers to catch the mistakes. It turns a gamble into an engineering problem.
Jane: And that's the real takeaway. This paper isn't just a warning. It's a guide. It's a way to make sure that when you delegate your negotiation to an AI, you're getting the deal you wanted, not the deal the AI's training data decided you should get.
Tom: It's a fantastic piece of research, and it's going to be a reference point for anyone working in this space. So, with that, we're going to say goodbye to "When LLM Agents Negotiate." It's been a real eye-opener.
Jane: It really has. Thanks for joining us, everyone. We'll be back soon with another paper that's shaping the future of AI. Until then, take care.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization