Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities

arXiv:2507.21790 · econ.EM, cs.AI · Submitted 2025-07-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities".

Jane: The paper was written by Nova, G., van Cranenburgh, S. and Hess, S. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 2: Tom: So, we've seen the title and the general scope, but what exactly did this paper find? The authors ran a massive experiment involving twelve versions of seven major LLM families—OpenAI, Anthropic, Google, DeepSeek—and tested them across five different configurations.

Jane: It’s clear that the paper summarized its findings by detailing how these models performed across three specific variables: the modeling goal (suggest versus suggest and estimate), the prompting strategy (Zero-Shot versus Chain-of-Thoughts), and whether they had full access to a dataset or just a data dictionary summary.

Lu: What stood out to me in the summary is that proprietary LLMs, like those from OpenAI and Anthropic, managed to generate valid and behaviorally sound utility specifications when guided by structured prompts, whereas open-weight models struggled significantly with meaningful outputs.

Meng: That’s a practical hurdle we need to address. The finding that constrained input—just the data dictionary—sometimes enhanced internal reasoning capabilities suggests that LLMs might benefit from being less distracted by raw data and more focused on the task itself.

Lalam: It's interesting how the paper implies that limiting raw data access can actually improve performance, suggesting a shift in how we design our AI agents to be more effective at conceptual work rather than just processing massive amounts of input.

Paper discussion segment 3: Tom: The authors identified a few key improvements or strengths they found, and this is where the results really shine. They noticed that while LLMs were capable of generating specifications, estimation was a whole other story.

Jane: It’s not just about suggesting the model; it's about executing it. The paper highlighted that only GPT-o3 was uniquely capable of correctly estimating its own generated specifications by running self-generated code in an agentic setting, which is a major milestone for end-to-end automation.

Lu: That capability of self-generating and executing code is huge for the future of AI in science. It means we are moving past just text output and into actual computational ability, which is a massive leap forward in the evolution of these tools.

Meng: From an engineering standpoint, that "agentic setting" where GPT-o3 could run code is what makes it reliable enough to be taken seriously in the modeling workflow, unlike the other non-agentic models that just hallucinated or misestimated everything else.

Lalam: The paper suggests that these advancements in tool use and structured reasoning are crucial for us to see AI not as a static calculator but as a dynamic partner that can improve our intellectual processes.

Paper discussion segment 4: Tom: We’ve looked at the general findings, but let's get into how the prompting strategy matters. The paper really shows that structure is key to quality.

Jane: It seems like the authors found that when using a Zero-Shot prompt, you get a wide variety of suggestions, but they consistently showed that applying Chain-of-Thoughts prompting led to significantly higher quality and more behaviourally plausible model specifications across all LLMs.

Lu: The Chain-of-Thoughts approach forces the LLM to perform descriptive analysis and step through the modeling workflow logically before proposing a utility function, which is exactly how an expert would approach it.

Meng: That structured guidance helps mitigate the risk of hallucination. By forcing the AI to follow a specific, multi-step reasoning process, we increase its reliability when trying to solve complex problems like MNL specification.

Lalam: The consistent improvement in quality from CoT prompts suggests that we are learning how to effectively teach these tools the logic and methodology of our specific fields, guiding their evolution toward better practices.

Conclusion: Tom: So, we’ve covered the results and the methodology, but let's wrap up what this all means for a final time "Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities."

Jane: The paper concludes that while LLMs show incredible potential in assisting with utility specification, they are currently support tools, not autonomous agents, and requires expert oversight.

Lu: I hope this shows us the pathway toward a future where AI and human expertise can work together seamlessly on complex scientific problems.

Meng: My final thought is that we need to keep pushing the boundaries of prompt engineering to make these systems as reliable as possible for real-world deployment.

Lalam: And my hope is that this research helps us build hybrid workflows where the AI handles the heavy lifting of suggestion while human expertise provides the necessary validation, leading to a more efficient and informed future culture.

Tom: That's a powerful way to end, guys. Thank you all for breaking down this complex paper with us today!

Jane: Join us next time for our deep dive into another fascinating arXiv preprint.

Nova, G., van Cranenburgh, S., Hess, S.

econ.EM, cs.AI

Submitted: 2025-07-29

Updated: 2026-08-25

Code: https://github.com/JonathanChavezTamales/llm-leaderboard

Importance score: 89/100

The gist: The field of choice modeling, which analyzes how individuals make decisions across multiple options, has historically relied on complex econometric frameworks.

Key concepts

Chain-of-Thoughts (CoT)
A prompting strategy where the AI is forced to perform descriptive analysis and step through a modeling workflow logically. This structured guidance helps mitigate hallucination and increases reliability when solving complex problems.
Zero-Shot Prompt
A prompting strategy where the AI is given a task without specific examples or detailed instructions. While it generates a wide variety of suggestions, the episode notes that Chain-of-Thoughts generally leads to higher quality results.
Utility Specification
The process of defining the mathematical model for consumer choice. The paper found that proprietary LLMs are capable of generating valid and behaviorally sound specifications when guided by structured prompts.

Terminology

Summary

The field of choice modeling, which analyzes how individuals make decisions across multiple options, has historically relied on complex econometric frameworks. This paper explores the emerging intersection between these established methodologies and modern Large Language Models (LLMs), arguing that LLMs can serve as powerful assistants in various stages of the choice modeling pipeline—from initial data preparation and specification to model inference and interpretation. By detailing advanced prompting strategies, the research aims to illuminate how large language models can assist choice modelling, thereby expanding the accessibility and sophistication of behavioral analysis for practitioners.

Foundations of Choice Modeling in a Machine Learning Context

Traditional choice models, such as those developed by Ortelli et al. (2021) and Rodrigues et al. (2020), require rigorous specification of utility functions to accurately predict consumer behavior. The process involves defining the underlying determinants that influence decision-making, a task that can be highly subjective or computationally intensive. The paper addresses the challenge posed by integrating these structured econometric models with the flexible, high-dimensional nature of machine learning outputs. It emphasizes that while LLMs are adept at pattern recognition, their application in choice modeling requires careful methodological guidance to ensure that the model building, inference and interpretation remain statistically sound (Rodrigues et al., 2024).

LLM Capabilities for Model Specification and Tool Use

A significant focus is placed on leveraging the LLM's ability to act as a computational agent. The paper draws parallels with advancements like Toolformer (Schick et al., 2023), suggesting that LLMs can be taught to use external tools—such as statistical packages or specialized choice modeling libraries—to perform necessary calculations. This capability moves beyond simple text generation, allowing the model to teach themselves to use tools for quantitative tasks. Furthermore, techniques such as prompt programming (Reynolds et al., 2021) are utilized to guide the LLM through multi-step decision processes, ensuring that the model systematically identifies relevant variables and potential functional forms for utility specification.

Advanced Reasoning Strategies for Complex Inference

To overcome the limitations of single-shot prompting, the paper details several advanced reasoning techniques that enhance an LLM's capacity to handle complex choice scenarios. These strategies mimic human expert reasoning processes:

  • Chain-of-Thought (CoT) Prompting: As demonstrated by Wei et al. (2022), CoT prompting elicits step-by-step reasoning, allowing the model to break down a complex utility function specification into manageable, logical components.

  • Least-to-Most Prompting: This method (Zhou et al., 2022) enables complex reasoning by structuring the problem as a series of increasingly difficult subproblems, which is crucial when dealing with multi-faceted choice sets.

  • ReAct Framework: Yao et al. (2023) propose combining reasoning and acting, allowing the LLM not only to reason about the model structure but also to execute simulated actions or queries against a defined knowledge base, simulating a robust analytical workflow.

Benchmarking and Future Directions for Adoption

The paper stresses that successful adoption requires rigorous validation. It notes that while general benchmarks like MMLU-pro (Wang et al., 2024) test broad knowledge, choice modeling demands specialized evaluation. The authors suggest the need for dedicated benchmarks, such as those emerging from the Gpqa benchmark (Rein et al., 2024), to test domain-specific reasoning. Ultimately, the utility of LLMs in this field depends on their ability to perform reliable tool learning (Qu et al., 2025) and maintain transparency in their decision pathways, ensuring that the insights provided are both novel and statistically justifiable.

Improvements for AI systems

(Initial Assessment: The provided bibliography indicates a convergence point between advanced Large Language Model capabilities (LLMs) and highly specialized, mathematically rigorous domains, specifically Choice Modeling and Econometrics. The primary gap is the lack of verifiable, domain-constrained reasoning within generative AI systems.)

Based on the synthesis of these sources—particularly the intersection of LLM architecture (Vaswani et al., Radford et al.), advanced prompting/reasoning techniques (Wei et al., Zhou et al., Yao et al.), and structured domain knowledge (Ortelli et al., Rodrigues & Pereira, Van Cranenburgh & Wang)—I propose three critical, interconnected improvements to the AI system architecture.


The Improvement:

We must move beyond general-purpose API calling (Schick et al., OpenAI) and implement a specialized, structured inference layer that sits between the LLM's prompt interpretation and the execution environment. This engine is designed to automatically translate high-level, ambiguous research questions (Why do people choose Option A over Option B?) into mathematically rigorous, executable econometric specifications (e.g., discrete choice models like Mixed Logit or Nested Logit).

How it Works:

  1. Input: The user provides raw data and a qualitative research goal (e.g., Estimate the impact of travel time and income on mode choice).

  2. LLM Role (Conceptualization): The LLM analyzes the input, identifies potential utility components, and proposes multiple plausible model structures (drawing from knowledge bases compiled from sources like Ortelli et al.).

  3. VCSE Role (Verification & Specification): The engine acts as a formal constraint checker. It takes the proposed model structure and verifies its statistical assumptions against established econometric theory. It then generates not just Python/R code, but a structured, LaTeX-formatted mathematical specification (Model Spec).

  4. Execution: The Model Spec is executed in a sandboxed environment using specialized statistical libraries (e.g., mlogit or dedicated R packages), ensuring that the output coefficients are statistically sound and interpretable within the context of choice theory.

What the Improved AI System Can Do:

It can function as an AI-powered econometric consultant. Instead of merely describing concepts, it specifies and executes complex statistical models—automatically handling tasks like determining appropriate functional forms, identifying necessary auxiliary variables (e.g., incorporating time or context), and generating fully documented, publication-ready model outputs with associated confidence intervals and diagnostic checks. This drastically reduces the need for specialized human expertise in model specification.

The system is then scored not on its answer, but on the fidelity of its execution path: did it select the correct underlying mathematical framework? Did it correctly identify and warn about necessary assumptions?

Sources

Related papers