Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games".
Jane: The paper was written by Eilam Shapira, Omer Madmon, Roi Reichart and Moshe Tennenholtz from Technion - Israel Institute of Technology.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: Welcome back to the show, everybody. Today we’re cracking open a paper that’s been making the rounds, and the title alone is a bit of a challenge to the whole field. It’s called “Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games.”
Jane: And Tom, I have to say, just reading that title out loud gets me excited. For years, if you wanted to know how people make economic decisions, you had to build a lab, recruit participants, and run experiments. This paper asks whether a large language model can just… generate that data for you.
Tom: Exactly. And the authors here are from the Technion in Israel — Eilam Shapira, Omer Madmon, Roi Reichart, and Moshe Tennenholtz. That’s a team with serious credentials in both AI and game theory.
Jane: Right, and that combination matters. This isn’t just a computer science trick. They’re working with a specific economic setup called a persuasion game. Think of it like a hotel trying to convince you to book a room by showing you a review.
Tom: So the “expert” — the hotel — has private information about the hotel’s true quality. They can choose which review to show you. The decision-maker, that’s you, has to decide whether to trust the review and book the hotel, or stay home.
Jane: And the kicker is, this happens over and over again, round after round. So the decision-maker can learn from past experience. Did the hotel lie to me last time? Should I punish them by not booking now?
Tom: Right. And the paper’s central question is a big one. Instead of collecting data from hundreds of human players to train a model to predict human choices, can we just have an LLM play the game thousands of times and use that synthetic data instead?
Jane: And the answer they found is a pretty strong “yes.” In fact, for a large enough sample, models trained on LLM-generated data actually beat models trained on real human data.
Tom: That’s a wild result. It suggests that the LLM, in this specific game, is capturing something fundamental about how people actually behave in these strategic situations. It’s not just mimicking language; it’s mimicking strategy.
Jane: It’s a bold claim, but the paper backs it up with a lot of experiments. And that’s what we’re going to dig into next. We’ll talk about how they actually set up the data collection and what the main results were.
Tom: Stay with us.
Summary of the Paper: Jane: So, Tom, we’ve set the stage. Let’s talk about what the paper actually did. The setup is a language-based persuasion game, and they had a dataset of over seventy-one thousand decisions from two hundred ten real human players.
Tom: Right, that’s the human baseline. Then they took five different LLMs — we’re talking Qwen, Llama, Gemini, Chat-Bison — and had them play the exact same game against the same six expert strategies.
Jane: And to make the synthetic data more varied, they gave each LLM player a “persona.” Some were told to be optimistic, some pessimistic, some cared about the price, others about the location. It’s a clever trick to get more diverse behavior out of the model.
Tom: And the results were striking. The best LLM, Qwen-two-72B, generated data that, when used to train a simple LSTM predictor, achieved about seventy-nine point one percent accuracy on human choices. That’s better than the seventy-seven point seven percent you get from training on all one hundred ten human players.
Lu: I want to jump in here, Jane. That’s the headline, but the more interesting part to me is *why* it works. The paper shows it’s not just about the sentiment of the review. They built a baseline that only looked at the text’s sentiment, and it did much worse.
Meng: Right, Lu. The LLM players are doing something more than just reading the review. They’re reacting to the history of the game. They’re building trust, they’re punishing the expert for lying. That’s what makes the data so rich.
Jane: Exactly. They even showed that if you restrict the LLM’s memory to just one or two rounds of history, the quality of the generated data drops significantly. The full history is crucial.
Tom: And that’s the key insight, isn’t it? This isn’t a sentiment analysis task. It’s a strategic decision-making task. The LLM, by playing the whole game, is learning the dynamics of trust and deception, and that’s what makes its predictions so good.
Lu: And that’s why the paper’s title is so provocative. It’s not just saying LLMs can predict choices. It’s saying they can generate the *training data* for those predictors, which is a much more powerful claim.
Meng: And a much more practical one. For a company like Booking.com, getting thousands of human decisions is expensive and slow. Generating them from an LLM is cheap and fast.
Jane: But it’s not a free lunch. We saw that the models trained purely on LLM data were less calibrated. Their confidence levels didn’t match reality as well as models trained on human data.
Tom: So you get higher accuracy but worse calibration. That’s a trade-off we’ll have to dig into. But first, let’s talk about the clever ways they tried to get the best of both worlds.
Improvements Suggested by the Paper: Tom: So we’ve established that LLM data is powerful but has a calibration problem. What’s the fix? The paper suggests a few smart ways to combine the two worlds.
Jane: The simplest one is just mixing the datasets. You take your one hundred ten human players and you add a bunch of LLM players to the training set. And that works. Accuracy goes up, and the calibration error drops significantly.
Meng: That’s the easy win. But they went further. They explored fine-tuning the LLM itself on the human data before using it to generate more data. And that’s where things get really interesting.
Tom: Right. They took Llama-three-8B, which was actually their weakest data generator, and fine-tuned it on the one hundred ten human players. And suddenly, the data it generated was good enough to train a predictor that hit eighty point one percent accuracy. That’s better than anything else they had.
Lu: That’s a huge jump. It shows that a small amount of high-quality human data can be used to “specialize” a general-purpose LLM into a much better simulator of human behavior. It’s like giving the model a crash course in human psychology.
Jane: And then they had this clever idea called DUAL — Double Use of human data for Augmented Learning. You use the human data to fine-tune the LLM, and then you *also* use that same human data to train the final predictor, alongside the new synthetic data.
Meng: And the result? The accuracy stays at that high eighty point one percent, but the calibration error gets cut almost in half, from zero point one five down to zero point zero eight. So you get the best of both worlds.
Tom: So the human data is doing double duty. It’s teaching the LLM how to behave, and it’s also grounding the final predictor in reality.
Lu: It’s a really elegant solution. It acknowledges that synthetic data is great for capturing the broad patterns, but you still need a bit of real human data to anchor the model’s confidence.
Jane: And this isn’t just a theoretical exercise. The paper shows this works across different prediction models — LSTM, Mamba, Transformer, even XGBoost. And they also tested it in a completely different persuasion game framework called GLEE, and the results held up.
Meng: That’s the part that gets me excited as an engineer. It’s not a one-off trick. It’s a robust methodology that could be applied to any situation where you need to predict human behavior in a strategic setting.
Tom: So we have a path forward. Use LLMs to generate a ton of data, use a little bit of human data to fine-tune the generator, and then mix everything together for a final, well-calibrated predictor.
Conclusion: Tom: Alright, let’s wrap this up. We’ve been talking about “Can LLMs Replace Economic Choice Prediction Labs?” and I think we’ve got a clear picture now.
Jane: We have. The paper shows that LLMs can absolutely generate training data for human choice prediction, and they can even do it better than using human data alone. But the real magic happens when you combine the two.
Tom: The DUAL approach — using human data to fine-tune the LLM and then mixing it back into the training set — gives you the highest accuracy and the best calibration. It’s a blueprint for how to do this right.
Lu: And the deeper lesson is about what makes human decisions predictable. It’s not just the words in front of you. It’s the history of the interaction. The LLMs that captured that history best were the ones that generated the most useful data.
Meng: From a practical standpoint, this could save companies and researchers a fortune. Instead of running massive human experiments, you can run a few small ones to fine-tune an LLM, and then scale up synthetic data generation for free.
Jane: It’s a powerful tool, but we should remember the ethical side. Predicting human choices more accurately could be used for good, like better user interfaces, or for manipulation. The authors are right to call for responsible use.
Tom: Absolutely. But as a piece of research, this is a landmark. It shows a new way to do behavioral economics, one that’s faster, cheaper, and potentially even more accurate.
Jane: So we’ll say goodbye to this paper with a sense of excitement for what comes next. Can this approach work in other domains? We’ll have to see.
Tom: Thanks for listening, everyone. We’ll be back with another paper soon.
Eilam Shapira, Omer Madmon, Roi Reichart, Moshe Tennenholtz
Technion - Israel Institute of Technology
cs.LG, cs.AI, cs.CL, cs.GT, cs.HC
Submitted: 2025-11-20
Updated: 2026-08-18
Code: https://github.com/ilamshapira/HumanChoicePrediction
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: This paper investigates whether Large Language Models (LLMs) can generate training data for human choice prediction in economic settings, specifically within language-based persuasion games.
Key concepts
- Persuasion Game
- An economic setup where an 'expert' (like a hotel) has private information about quality and chooses which review to show. The decision-maker must decide whether to trust the review and book, or stay home, over multiple rounds.
- LLM-Generated Data
- Synthetic data created when an LLM plays a strategic game thousands of times. The paper suggests this data can be used to train models that predict human behavior, potentially replacing expensive human experiments.
- DUAL Approach
- A methodology combining the strengths of both data sources. It involves using limited real human data to fine-tune an LLM (the generator), and then mixing that synthetic output with the original human data for the final predictor.
Terminology
Summary
This paper investigates whether Large Language Models (LLMs) can generate training data for human choice prediction in economic settings, specifically within language-based persuasion games. The authors demonstrate that models trained on LLM-generated data can effectively predict human behavior and even outperform models trained on actual human data. They also explore the dual role of LLMs as both data generators and predictors, introduce a strategy called Double Use of human data for Augmented Learning (DUAL),
and analyze the importance of interaction history in predicting human decisions.
The paper begins by describing the language-based persuasion game, which involves an expert (sender) and a decision-maker (DM) interacting over multiple rounds. The expert has private information about a hotel's quality and sends a textual review to the DM, who decides whether to go to the hotel
or stay home.
The expert always benefits from the DM opting in, while the DM only benefits if the hotel is of high quality. The task is to predict human DM behavior against a set of expert strategies.
The authors collected a human dataset of 71,579 decisions from 210 distinct human players interacting with six expert strategies. They then created LLM-generated datasets by replicating the process with five different LLMs (Chat-Bison, Gemini-1.5, Qwen-2-72B, Llama-3-70B, and Llama-3-8B) and used persona diversification
to improve sample complexity.
Key findings include:
-
LLM-generated data effectiveness:
For all types of prediction models, LLM-generated data outperforms human data in the human choice task for a large enough sample size.
The Qwen-2-72B LLM achieved the best performance. -
Combining data sources:
Enriching the human data with synthetic data significantly improved performance,
and a model trained on a mixture of LLM-based and human players outperforms a model trained on human data only. -
Calibration trade-off:
The ECE values of models trained using LLM data are higher than those of models trained on human data, indicating a degradation in calibration.
However, combining LLM-generated and human data results in the most calibrated model. -
Per-strategy analysis: The LLM-based approach
always outperforms the models trained on 32 human players
and sometimes outperforms models trained on 110 human players. The only exception where the sentiment baseline outperforms is against the Honest strategy. -
Fine-tuning and DUAL: Fine-tuning an LLM on human data improves its data generation quality, and the DUAL approach (reusing human data for both fine-tuning and training the predictor)
nearly halves calibration error (ECE), from 0.15 to 0.08
while maintaining accuracy. -
Importance of history:
Interaction history plays a crucial role in predicting and explaining human behavior.
The authors found that "similarity to human actions when grouped by history is positively correlated with the predictive power of models trained on this data, while similarity to human actions when grouped by sentiment actually correlates negatively with it." Furthermore, restricting the LLM's visible history to one or two turns leads to a noticeable drop in predictive accuracy. -
Generalization: The authors provide preliminary evidence that the effectiveness of LLM-generated data extends to other persuasion game frameworks, such as GLEE, where models trained on LLM-vs-LLM games outperformed models trained on limited human data.
The paper concludes that training a choice prediction model on a dataset containing no human choice data at all can even outperform the same model trained on an actual human-generated dataset, given enough generated data points.
The authors also discuss limitations, including the focus on a specific class of games and the calibration issue, and highlight ethical considerations regarding the potential for misuse of human choice prediction.
Improvements for AI systems
Based on the paper, here are specific improvements that can be made to AI systems, along with the resulting capabilities.
Improvement: Implement a two-stage pipeline where a fine-tuned LLM is used as a data generator to create synthetic training data for a separate, lightweight prediction model (e.g., LSTM).
-
Specific Action: Instead of using an off-the-shelf LLM, fine-tune a smaller model (like Llama-3-8B) on a small set of human interaction data. Then, use this fine-tuned model to generate a large volume of synthetic interaction data. Finally, train a simple LSTM classifier on this synthetic data.
-
What the Improved System Can Do: Achieve higher prediction accuracy (80.1% vs. 79.1% for off-the-shelf LLM data) than using human data alone (77.7%) or data from larger, non-fine-tuned models. This makes it cost-effective and scalable, as generating synthetic data is cheaper than collecting more human data.
Abstract
Human choice prediction in economic contexts is crucial for applications in marketing, finance, public policy, and more. This task, however, is often constrained by the difficulties in acquiring human choice data. With most experimental economics studies focusing on simple choice settings, the AI community has explored whether LLMs can substitute for humans in these predictions and examined more complex experimental economics settings. However, a key question remains: can LLMs generate training data for human choice prediction? We explore this in language-based persuasion games, a complex economic setting involving natural language in strategic interactions. Our experiments show that models trained on LLM-generated data can effectively predict human behavior in these games and even outperform models trained on actual human data. Beyond data generation, we investigate the dual role of LLMs as both data generators and predictors, introducing a comprehensive empirical study on the effectiveness of utilizing LLMs for data generation, human choice prediction, or both. We then utilize our choice prediction framework to analyze how strategic factors shape decision-making, showing that interaction history (rather than linguistic sentiment alone) plays a key role in predicting human decision-making in repeated interactions. Particularly, when LLMs capture history-dependent decision patterns similarly to humans, their predictive success improves substantially. Finally, we demonstrate the robustness of our findings across alternative persuasion-game settings, highlighting the broader potential of using LLM-generated data to model human decision-making.
Sources
- Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses
- PaLM 2 Technical Report
- Generating Synthetic Documents for Cross-Encoder Re-Rankers: A Comparative Study of ChatGPT and Human Experts
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
- Information Design With Large Language Models
- A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios
- Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- GPT-4 Technical Report
- Do LLM Agents Have Regret? A Case Study in Online Learning and Games
- Predicting human decisions with behavioral theories and machine learning
- How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
- GLEE: A Unified Framework and Benchmark for Language-based Economic Environments
- The Curse of Recursion: Training on Generated Data Makes Models Forget
- Simulating Human Strategic Behavior: Comparing Single and Multi-agent LLMs
- Putting GPT-3's Creativity to the (Alternative Uses) Test
- Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks