Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games
summary
The gist
This paper investigates whether Large Language Models (LLMs) can generate training data for human choice prediction in economic settings, specifically within language-based persuasion games.
In short
The episode discusses a paper exploring if Large Language Models (LLMs) can replace human participants in predicting economic choices. Hosts find that LLM-generated data is highly effective, but the best results come from a 'DUAL' approach: using limited human data to fine-tune the LLM, and then mixing that synthetic data with the real data for prediction.
Key concepts
- Persuasion Game
- An economic setup where an 'expert' (like a hotel) has private information about quality and chooses which review to show. The decision-maker must decide whether to trust the review and book, or stay home, over multiple rounds.
- LLM-Generated Data
- Synthetic data created when an LLM plays a strategic game thousands of times. The paper suggests this data can be used to train models that predict human behavior, potentially replacing expensive human experiments.
- DUAL Approach
- A methodology combining the strengths of both data sources. It involves using limited real human data to fine-tune an LLM (the generator), and then mixing that synthetic output with the original human data for the final predictor.
Terminology used across episodes
This episode discusses
- Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games · Paper Radio
- Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
- PaLM 2 Technical Report
- Generating Synthetic Documents for Cross-Encoder Re-Rankers: A Comparative Study of ChatGPT and Human Experts
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
- Information Design With Large Language Models
- A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios
- Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- GPT-4 Technical Report
- Do LLM Agents Have Regret? A Case Study in Online Learning and Games
- Predicting human decisions with behavioral theories and machine learning
- How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
- GLEE: A Unified Framework and Benchmark for Language-based Economic Environments
- The Curse of Recursion: Training on Generated Data Makes Models Forget
- Simulating Human Strategic Behavior: Comparing Single and Multi-agent LLMs
- Putting GPT-3's Creativity to the (Alternative Uses) Test
- Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers
The paper
Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games · Read on arXiv
Eilam Shapira, Omer Madmon, Roi Reichart, Moshe Tennenholtz
Technion - Israel Institute of Technology
Human choice prediction in economic contexts is crucial for applications in marketing, finance, public policy, and more. This task, however, is often constrained by the difficulties in acquiring human choice data. With most experimental economics studies focusing on simple choice settings, the AI community has explored whether LLMs can substitute for humans in these predictions and examined more complex experimental economics settings. However, a key question remains: can LLMs generate training data for human choice prediction? We explore this in language-based persuasion games, a complex economic setting involving natural language in strategic interactions. Our experiments show that models trained on LLM-generated data can effectively predict human behavior in these games and even outperform models trained on actual human data. Beyond data generation, we investigate the dual role of LLMs as both data generators and predictors, introducing a comprehensive empirical study on the effectiveness of utilizing LLMs for data generation, human choice prediction, or both. We then utilize our choice prediction framework to analyze how strategic factors shape decision-making, showing that interaction history (rather than linguistic sentiment alone) plays a key role in predicting human decision-making in repeated interactions. Particularly, when LLMs capture history-dependent decision patterns similarly to humans, their predictive success improves substantially. Finally, we demonstrate the robustness of our findings across alternative persuasion-game settings, highlighting the broader potential of using LLM-generated data to model human decision-making.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games".
Jane: The paper was written by Eilam Shapira, Omer Madmon, Roi Reichart and Moshe Tennenholtz from Technion - Israel Institute of Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: Welcome back to the show, everybody. Today we’re cracking open a paper that’s been making the rounds, and the title alone is a bit of a challenge to the whole field. It’s called “Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games.”
Jane: And Tom, I have to say, just reading that title out loud gets me excited. For years, if you wanted to know how people make economic decisions, you had to build a lab, recruit participants, and run experiments. This paper asks whether a large language model can just… generate that data for you.
Tom: Exactly. And the authors here are from the Technion in Israel — Eilam Shapira, Omer Madmon, Roi Reichart, and Moshe Tennenholtz. That’s a team with serious credentials in both AI and game theory.
Jane: Right, and that combination matters. This isn’t just a computer science trick. They’re working with a specific economic setup called a persuasion game. Think of it like a hotel trying to convince you to book a room by showing you a review.
Tom: So the “expert” — the hotel — has private information about the hotel’s true quality. They can choose which review to show you. The decision-maker, that’s you, has to decide whether to trust the review and book the hotel, or stay home.
Jane: And the kicker is, this happens over and over again, round after round. So the decision-maker can learn from past experience. Did the hotel lie to me last time? Should I punish them by not booking now?
Tom: Right. And the paper’s central question is a big one. Instead of collecting data from hundreds of human players to train a model to predict human choices, can we just have an LLM play the game thousands of times and use that synthetic data instead?
Jane: And the answer they found is a pretty strong “yes.” In fact, for a large enough sample, models trained on LLM-generated data actually beat models trained on real human data.
Tom: That’s a wild result. It suggests that the LLM, in this specific game, is capturing something fundamental about how people actually behave in these strategic situations. It’s not just mimicking language; it’s mimicking strategy.
Jane: It’s a bold claim, but the paper backs it up with a lot of experiments. And that’s what we’re going to dig into next. We’ll talk about how they actually set up the data collection and what the main results were.
Tom: Stay with us.
Summary of the Paper: Jane: So, Tom, we’ve set the stage. Let’s talk about what the paper actually did. The setup is a language-based persuasion game, and they had a dataset of over seventy-one thousand decisions from two hundred ten real human players.
Tom: Right, that’s the human baseline. Then they took five different LLMs — we’re talking Qwen, Llama, Gemini, Chat-Bison — and had them play the exact same game against the same six expert strategies.
Jane: And to make the synthetic data more varied, they gave each LLM player a “persona.” Some were told to be optimistic, some pessimistic, some cared about the price, others about the location. It’s a clever trick to get more diverse behavior out of the model.
Tom: And the results were striking. The best LLM, Qwen-two-72B, generated data that, when used to train a simple LSTM predictor, achieved about seventy-nine point one percent accuracy on human choices. That’s better than the seventy-seven point seven percent you get from training on all one hundred ten human players.
Lu: I want to jump in here, Jane. That’s the headline, but the more interesting part to me is *why* it works. The paper shows it’s not just about the sentiment of the review. They built a baseline that only looked at the text’s sentiment, and it did much worse.
Meng: Right, Lu. The LLM players are doing something more than just reading the review. They’re reacting to the history of the game. They’re building trust, they’re punishing the expert for lying. That’s what makes the data so rich.
Jane: Exactly. They even showed that if you restrict the LLM’s memory to just one or two rounds of history, the quality of the generated data drops significantly. The full history is crucial.
Tom: And that’s the key insight, isn’t it? This isn’t a sentiment analysis task. It’s a strategic decision-making task. The LLM, by playing the whole game, is learning the dynamics of trust and deception, and that’s what makes its predictions so good.
Lu: And that’s why the paper’s title is so provocative. It’s not just saying LLMs can predict choices. It’s saying they can generate the *training data* for those predictors, which is a much more powerful claim.
Meng: And a much more practical one. For a company like Booking.com, getting thousands of human decisions is expensive and slow. Generating them from an LLM is cheap and fast.
Jane: But it’s not a free lunch. We saw that the models trained purely on LLM data were less calibrated. Their confidence levels didn’t match reality as well as models trained on human data.
Tom: So you get higher accuracy but worse calibration. That’s a trade-off we’ll have to dig into. But first, let’s talk about the clever ways they tried to get the best of both worlds.
Improvements Suggested by the Paper: Tom: So we’ve established that LLM data is powerful but has a calibration problem. What’s the fix? The paper suggests a few smart ways to combine the two worlds.
Jane: The simplest one is just mixing the datasets. You take your one hundred ten human players and you add a bunch of LLM players to the training set. And that works. Accuracy goes up, and the calibration error drops significantly.
Meng: That’s the easy win. But they went further. They explored fine-tuning the LLM itself on the human data before using it to generate more data. And that’s where things get really interesting.
Tom: Right. They took Llama-three-8B, which was actually their weakest data generator, and fine-tuned it on the one hundred ten human players. And suddenly, the data it generated was good enough to train a predictor that hit eighty point one percent accuracy. That’s better than anything else they had.
Lu: That’s a huge jump. It shows that a small amount of high-quality human data can be used to “specialize” a general-purpose LLM into a much better simulator of human behavior. It’s like giving the model a crash course in human psychology.
Jane: And then they had this clever idea called DUAL — Double Use of human data for Augmented Learning. You use the human data to fine-tune the LLM, and then you *also* use that same human data to train the final predictor, alongside the new synthetic data.
Meng: And the result? The accuracy stays at that high eighty point one percent, but the calibration error gets cut almost in half, from zero point one five down to zero point zero eight. So you get the best of both worlds.
Tom: So the human data is doing double duty. It’s teaching the LLM how to behave, and it’s also grounding the final predictor in reality.
Lu: It’s a really elegant solution. It acknowledges that synthetic data is great for capturing the broad patterns, but you still need a bit of real human data to anchor the model’s confidence.
Jane: And this isn’t just a theoretical exercise. The paper shows this works across different prediction models — LSTM, Mamba, Transformer, even XGBoost. And they also tested it in a completely different persuasion game framework called GLEE, and the results held up.
Meng: That’s the part that gets me excited as an engineer. It’s not a one-off trick. It’s a robust methodology that could be applied to any situation where you need to predict human behavior in a strategic setting.
Tom: So we have a path forward. Use LLMs to generate a ton of data, use a little bit of human data to fine-tune the generator, and then mix everything together for a final, well-calibrated predictor.
Conclusion: Tom: Alright, let’s wrap this up. We’ve been talking about “Can LLMs Replace Economic Choice Prediction Labs?” and I think we’ve got a clear picture now.
Jane: We have. The paper shows that LLMs can absolutely generate training data for human choice prediction, and they can even do it better than using human data alone. But the real magic happens when you combine the two.
Tom: The DUAL approach — using human data to fine-tune the LLM and then mixing it back into the training set — gives you the highest accuracy and the best calibration. It’s a blueprint for how to do this right.
Lu: And the deeper lesson is about what makes human decisions predictable. It’s not just the words in front of you. It’s the history of the interaction. The LLMs that captured that history best were the ones that generated the most useful data.
Meng: From a practical standpoint, this could save companies and researchers a fortune. Instead of running massive human experiments, you can run a few small ones to fine-tune an LLM, and then scale up synthetic data generation for free.
Jane: It’s a powerful tool, but we should remember the ethical side. Predicting human choices more accurately could be used for good, like better user interfaces, or for manipulation. The authors are right to call for responsible use.
Tom: Absolutely. But as a piece of research, this is a landmark. It shows a new way to do behavioral economics, one that’s faster, cheaper, and potentially even more accurate.
Jane: So we’ll say goodbye to this paper with a sense of excitement for what comes next. Can this approach work in other domains? We’ll have to see.
Tom: Thanks for listening, everyone. We’ll be back with another paper soon.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization