Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities
summary
The gist
The field of choice modeling, which analyzes how individuals make decisions across multiple options, has historically relied on complex econometric frameworks.
In short
The episode discusses a paper exploring how large language models (LLMs) can assist in choice modeling. The hosts detail experiments testing various LLMs across different prompting strategies and data access levels, concluding that while LLMs are powerful for suggesting utility specifications, they remain support tools requiring expert oversight.
Key concepts
- Chain-of-Thoughts (CoT)
- A prompting strategy where the AI is forced to perform descriptive analysis and step through a modeling workflow logically. This structured guidance helps mitigate hallucination and increases reliability when solving complex problems.
- Zero-Shot Prompt
- A prompting strategy where the AI is given a task without specific examples or detailed instructions. While it generates a wide variety of suggestions, the episode notes that Chain-of-Thoughts generally leads to higher quality results.
- Utility Specification
- The process of defining the mathematical model for consumer choice. The paper found that proprietary LLMs are capable of generating valid and behaviorally sound specifications when guided by structured prompts.
Terminology used across episodes
This episode discusses
- Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities · Paper Radio
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3 Technical Report
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Gemini: A Family of Highly Capable Multimodal Models
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- GPT-4 Technical Report
- Jamba-1.5: Hybrid Transformer-Mamba Models at Scale
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- HelpSteer2-Preference: Complementing Ratings with Preferences
- Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Phi-4 Technical Report
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
- Phi-4-reasoning Technical Report
- Qwen2 Technical Report
- Qwen2.5-Coder Technical Report
- YaRN: Efficient Context Window Extension of Large Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
The paper
Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities · Read on arXiv
Nova, G., van Cranenburgh, S., Hess, S.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities".
Jane: The paper was written by Nova, G., van Cranenburgh, S. and Hess, S. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 2: Tom: So, we've seen the title and the general scope, but what exactly did this paper find? The authors ran a massive experiment involving twelve versions of seven major LLM families—OpenAI, Anthropic, Google, DeepSeek—and tested them across five different configurations.
Jane: It’s clear that the paper summarized its findings by detailing how these models performed across three specific variables: the modeling goal (suggest versus suggest and estimate), the prompting strategy (Zero-Shot versus Chain-of-Thoughts), and whether they had full access to a dataset or just a data dictionary summary.
Lu: What stood out to me in the summary is that proprietary LLMs, like those from OpenAI and Anthropic, managed to generate valid and behaviorally sound utility specifications when guided by structured prompts, whereas open-weight models struggled significantly with meaningful outputs.
Meng: That’s a practical hurdle we need to address. The finding that constrained input—just the data dictionary—sometimes enhanced internal reasoning capabilities suggests that LLMs might benefit from being less distracted by raw data and more focused on the task itself.
Lalam: It's interesting how the paper implies that limiting raw data access can actually improve performance, suggesting a shift in how we design our AI agents to be more effective at conceptual work rather than just processing massive amounts of input.
Paper discussion segment 3: Tom: The authors identified a few key improvements or strengths they found, and this is where the results really shine. They noticed that while LLMs were capable of generating specifications, estimation was a whole other story.
Jane: It’s not just about suggesting the model; it's about executing it. The paper highlighted that only GPT-o3 was uniquely capable of correctly estimating its own generated specifications by running self-generated code in an agentic setting, which is a major milestone for end-to-end automation.
Lu: That capability of self-generating and executing code is huge for the future of AI in science. It means we are moving past just text output and into actual computational ability, which is a massive leap forward in the evolution of these tools.
Meng: From an engineering standpoint, that "agentic setting" where GPT-o3 could run code is what makes it reliable enough to be taken seriously in the modeling workflow, unlike the other non-agentic models that just hallucinated or misestimated everything else.
Lalam: The paper suggests that these advancements in tool use and structured reasoning are crucial for us to see AI not as a static calculator but as a dynamic partner that can improve our intellectual processes.
Paper discussion segment 4: Tom: We’ve looked at the general findings, but let's get into how the prompting strategy matters. The paper really shows that structure is key to quality.
Jane: It seems like the authors found that when using a Zero-Shot prompt, you get a wide variety of suggestions, but they consistently showed that applying Chain-of-Thoughts prompting led to significantly higher quality and more behaviourally plausible model specifications across all LLMs.
Lu: The Chain-of-Thoughts approach forces the LLM to perform descriptive analysis and step through the modeling workflow logically before proposing a utility function, which is exactly how an expert would approach it.
Meng: That structured guidance helps mitigate the risk of hallucination. By forcing the AI to follow a specific, multi-step reasoning process, we increase its reliability when trying to solve complex problems like MNL specification.
Lalam: The consistent improvement in quality from CoT prompts suggests that we are learning how to effectively teach these tools the logic and methodology of our specific fields, guiding their evolution toward better practices.
Conclusion: Tom: So, we’ve covered the results and the methodology, but let's wrap up what this all means for a final time "Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities."
Jane: The paper concludes that while LLMs show incredible potential in assisting with utility specification, they are currently support tools, not autonomous agents, and requires expert oversight.
Lu: I hope this shows us the pathway toward a future where AI and human expertise can work together seamlessly on complex scientific problems.
Meng: My final thought is that we need to keep pushing the boundaries of prompt engineering to make these systems as reliable as possible for real-world deployment.
Lalam: And my hope is that this research helps us build hybrid workflows where the AI handles the heavy lifting of suggestion while human expertise provides the necessary validation, leading to a more efficient and informed future culture.
Tom: That's a powerful way to end, guys. Thank you all for breaking down this complex paper with us today!
Jane: Join us next time for our deep dive into another fascinating arXiv preprint.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization