Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities

summary

Video file (mp4)

The gist

The field of choice modeling, which analyzes how individuals make decisions across multiple options, has historically relied on complex econometric frameworks.

In short

The episode discusses a paper exploring how large language models (LLMs) can assist in choice modeling. The hosts detail experiments testing various LLMs across different prompting strategies and data access levels, concluding that while LLMs are powerful for suggesting utility specifications, they remain support tools requiring expert oversight.

Key concepts

Chain-of-Thoughts (CoT)
A prompting strategy where the AI is forced to perform descriptive analysis and step through a modeling workflow logically. This structured guidance helps mitigate hallucination and increases reliability when solving complex problems.
Zero-Shot Prompt
A prompting strategy where the AI is given a task without specific examples or detailed instructions. While it generates a wide variety of suggestions, the episode notes that Chain-of-Thoughts generally leads to higher quality results.
Utility Specification
The process of defining the mathematical model for consumer choice. The paper found that proprietary LLMs are capable of generating valid and behaviorally sound specifications when guided by structured prompts.

Terminology used across episodes

This episode discusses

The paper

Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities · Read on arXiv

Nova, G., van Cranenburgh, S., Hess, S.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities".

Jane: The paper was written by Nova, G., van Cranenburgh, S. and Hess, S. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 2: Tom: So, we've seen the title and the general scope, but what exactly did this paper find? The authors ran a massive experiment involving twelve versions of seven major LLM families—OpenAI, Anthropic, Google, DeepSeek—and tested them across five different configurations.

Jane: It’s clear that the paper summarized its findings by detailing how these models performed across three specific variables: the modeling goal (suggest versus suggest and estimate), the prompting strategy (Zero-Shot versus Chain-of-Thoughts), and whether they had full access to a dataset or just a data dictionary summary.

Lu: What stood out to me in the summary is that proprietary LLMs, like those from OpenAI and Anthropic, managed to generate valid and behaviorally sound utility specifications when guided by structured prompts, whereas open-weight models struggled significantly with meaningful outputs.

Meng: That’s a practical hurdle we need to address. The finding that constrained input—just the data dictionary—sometimes enhanced internal reasoning capabilities suggests that LLMs might benefit from being less distracted by raw data and more focused on the task itself.

Lalam: It's interesting how the paper implies that limiting raw data access can actually improve performance, suggesting a shift in how we design our AI agents to be more effective at conceptual work rather than just processing massive amounts of input.

Paper discussion segment 3: Tom: The authors identified a few key improvements or strengths they found, and this is where the results really shine. They noticed that while LLMs were capable of generating specifications, estimation was a whole other story.

Jane: It’s not just about suggesting the model; it's about executing it. The paper highlighted that only GPT-o3 was uniquely capable of correctly estimating its own generated specifications by running self-generated code in an agentic setting, which is a major milestone for end-to-end automation.

Lu: That capability of self-generating and executing code is huge for the future of AI in science. It means we are moving past just text output and into actual computational ability, which is a massive leap forward in the evolution of these tools.

Meng: From an engineering standpoint, that "agentic setting" where GPT-o3 could run code is what makes it reliable enough to be taken seriously in the modeling workflow, unlike the other non-agentic models that just hallucinated or misestimated everything else.

Lalam: The paper suggests that these advancements in tool use and structured reasoning are crucial for us to see AI not as a static calculator but as a dynamic partner that can improve our intellectual processes.

Paper discussion segment 4: Tom: We’ve looked at the general findings, but let's get into how the prompting strategy matters. The paper really shows that structure is key to quality.

Jane: It seems like the authors found that when using a Zero-Shot prompt, you get a wide variety of suggestions, but they consistently showed that applying Chain-of-Thoughts prompting led to significantly higher quality and more behaviourally plausible model specifications across all LLMs.

Lu: The Chain-of-Thoughts approach forces the LLM to perform descriptive analysis and step through the modeling workflow logically before proposing a utility function, which is exactly how an expert would approach it.

Meng: That structured guidance helps mitigate the risk of hallucination. By forcing the AI to follow a specific, multi-step reasoning process, we increase its reliability when trying to solve complex problems like MNL specification.

Lalam: The consistent improvement in quality from CoT prompts suggests that we are learning how to effectively teach these tools the logic and methodology of our specific fields, guiding their evolution toward better practices.

Conclusion: Tom: So, we’ve covered the results and the methodology, but let's wrap up what this all means for a final time "Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities."

Jane: The paper concludes that while LLMs show incredible potential in assisting with utility specification, they are currently support tools, not autonomous agents, and requires expert oversight.

Lu: I hope this shows us the pathway toward a future where AI and human expertise can work together seamlessly on complex scientific problems.

Meng: My final thought is that we need to keep pushing the boundaries of prompt engineering to make these systems as reliable as possible for real-world deployment.

Lalam: And my hope is that this research helps us build hybrid workflows where the AI handles the heavy lifting of suggestion while human expertise provides the necessary validation, leading to a more efficient and informed future culture.

Tom: That's a powerful way to end, guys. Thank you all for breaking down this complex paper with us today!

Jane: Join us next time for our deep dive into another fascinating arXiv preprint.

More episodes

← Home