Mining Causality: AI-Assisted Search for Instrumental Variables

arXiv:2409.14202 · econ.EM, stat.AP, stat.ME, stat.ML · Submitted 2024-09-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Mining Causality: AI-Assisted Search for Instrumental Variables".

Jane: This paper proposes a novel method for finding instrumental variables (IVs) by leveraging large language models (LLMs) to assist in the heuristic and creative process of causal inference.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We're looking at "Mining Causality: AI-Assisted Search for Instrumental Variables," and it’s clear the title tells you exactly what this paper is about—using artificial intelligence to help us hunt down those crucial instrumental variables. The authors are Sukjin Han and they’ve put forward a method that aims to make the search process much more systematic than what we typically see in causal inference research.

Jane: Right, Tom; it sounds like the focus isn't just on finding IVs, but specifically on tackling the difficult part of justifying those exclusion restrictions, which is where a lot of human creativity gets involved. The authors are using LLMs as a tool to simulate that creative decision-making process in a more structured way.

Lu: I think the authors are really pushing the idea that this isn't just about finding variables; it’s about how we interact with the search space itself, suggesting that LLMs can explore an exponentially faster rate and cover an extremely large search space than human researchers can manage alone.

Meng: From my side, I'm thinking about the practical application; if the method is so effective at navigating those complex scenarios, what does that mean for someone trying to build a rigorous study in a real-world setting? Does it reduce the time spent on literature review substantially?

Lalam: I feel like the authors are emphasizing that this approach goes beyond just finding variables; they are looking at how interacting with these AI tools can inspire completely new domains for novel IVs, which is a significant shift in how we think about discovery.

The paper's summary: Tom: So, the paper summarizes their core idea as using LLMs to search for new IVs through narratives and counterfactual reasoning, specifically arguing that multi-step and role-playing prompting strategies are effective for simulating how economic agents make decisions. They show how you can structure these prompts to test the relevance and exclusion assumptions systematically.

Jane: That means they’re not just asking the AI what IVs exist; they are asking it to *be* an agent in a scenario and figure out what factors would logically determine an outcome without directly affecting that outcome except through the treatment, which is a very concrete way to test those assumptions.

Lu: I find the multi-step approach really smart because it allows for subtask focus; you can first get a verbal description of relevance and exclusion, and then refine that set by asking the AI to select factors that are most likely independent, which makes sense for navigating complex decision logic.

Meng: That systematic breakdown sounds like it would be useful for structuring a research project; having a clear, repeatable sequence of prompts to guide the search rather than just throwing vague ideas at a model seems like a solid engineering approach.

Lalam: I think the summary highlights that this method is about using AI to navigate real-world scenarios rather than staying confined to academic texts, which is where I see the biggest cultural impact—it opens up new avenues for discovery outside of established literature.

The paper's improvements: Tom: Moving into how they improve the search, the paper focuses heavily on constructing specific prompt templates. They propose a very detailed structure for these prompts, like one asking an agent to identify factors that determine an outcome but don't directly affect it except through a treatment.

Jane: Those role-playing prompts are crucial because they align directly with where endogeneity often originates—in the agent’s decision itself; by framing it as a decision, you are simulating the endogenous process rather than just searching for abstract variables.

Lu: The authors detail specific templates, like Prompt one which asks for factors that determine an outcome but only through a treatment, and then Prompt two where the AI is tasked with choosing factors most likely to be unassociated with confounders based on what it just learned.

Meng: From an engineering viewpoint, those specific templates give us a clear blueprint for how to structure the interaction; it’s not just "ask the AI," but "here is exactly how you ask it to think about exclusion and independence," which makes the system much more controllable.

Lalam: The improvement here is really in moving from vague academic searching to having a repeatable, structured protocol that helps explore potential IVs through these specific conversational steps, which feels like a real advancement in the heuristic process.

Conclusion: Tom: To wrap up, the authors conclude that while the lists they generate aren't absolute truths, they serve as a valuable benchmark that inspires empirical researchers by providing strong starting points for their own work on finding instrumental variables. They are essentially positioning AI not just as a data-processing tool, but as something to assist in the creative and heuristic parts of causal inference.

Jane: So the big implication is that this approach helps us systematically search for exogeneity, treating AI as a partner in the creative discovery process rather than just a calculator or literature retriever. It suggests that researchers can gain multiple IV candidates by using this systematic search method.

Lu: I think the real impact is setting a new standard for how we use these models in causal inference; it shows that we can leverage their ability to explore large spaces and perform counterfactual reasoning to help build more robust identification strategies.

Meng: Practically, this means researchers could drastically cut down the initial manual literature review when trying to find candidate variables for complex designs like difference-in-differences or regression discontinuity. It’s about making the discovery phase much more efficient for empirical work.

Lalam: I think this paper paves a path toward using AI as a means to improve conventional research practices by focusing on that creative, heuristic side of causal inference, which is a huge step in how we approach complex identification problems.

Tom: That’s all the time for today. We’ve been talking about "Mining Causality: AI-Assisted Search for Instrumental Variables" and how this paper suggests a structured, multi-step prompting strategy can help researchers systematically search for IVs by simulating agent decision-making.

Jane: It's been fascinating to see how the authors use these LLMs to tackle the notoriously tricky task of justification through narrative and counterfactual exploration. We’ll be back next time with more insights into this field and other papers on arXiv.

School of Economics, University of Bristol

econ.EM, stat.AP, stat.ME, stat.ML

Submitted: 2024-09-21

Updated: 2026-09-30

Importance score: 86/100

The gist: This paper proposes a novel method for finding instrumental variables (IVs) by leveraging large language models (LLMs) to assist in the heuristic and creative process of causal inference.

Key concepts

Instrumental Variables (IVs)
These are external factors used in econometrics to estimate the causal effect of a variable when traditional methods fail due to endogeneity. They must be related to the treatment but not directly affect the outcome, making them crucial for isolating true cause and effect.
Endogeneity
This is a major problem in causal inference where the variable being studied is correlated with the error term in a regression model. This correlation makes it impossible to determine if observed effects are due to the treatment or some other unmeasured factor, thus biasing traditional statistical results.
Relevance (REL), Exclusion (EX), Independence (IND)
These are the three core assumptions an IV must satisfy: Relevance means the instrument must affect the treatment. Exclusion means it shouldn't directly affect the outcome except through the treatment. Independence means it should not be correlated with confounding factors, ensuring a clean causal path.

Terminology

Summary

This paper proposes a novel method for finding instrumental variables (IVs) by leveraging large language models (LLMs) to assist in the heuristic and creative process of causal inference. It argues that LLMs can significantly accelerate this search by exploring an extremely large search space and facilitating counterfactual reasoning, potentially allowing human researchers to discover new IVs more effectively than traditional methods.

The Core Problem and Proposed Solution

Endogeneity is a major obstacle in causal inference, and justifying exclusion restrictions—a key part of the IV method—is often a rhetorical process that relies on human creativity. The authors contend that LLMs are well-suited to assist in this search because they can dramatically accelerate this process and explore an extremely large search space, going beyond the narrow academic discourses on IVs. The goal is to use LLMs as sophisticated tools to simulate the endogenous decision-making processes of economic agents and navigate real-world scenarios rather than being anchored solely in academic texts.

The Multi-Step Prompting Strategy

The proposed method involves a two-step approach designed to systematically search for IVs that satisfy the three core assumptions: Relevance (REL), Exclusion (EX), and Independence (IND).

  1. In Step 1, the LLM is prompted to search for IVs satisfying a verbal description of REL and EX, often using role-playing prompts where the agent makes a decision in a hypothetical scenario.

  2. In Step 2, the LLM refines this set by selecting factors that satisfy IND (independence), often by querying it to choose variables most likely to be unassociated with [confounders].

The authors advocate for a multi-step approach because it allows for subtask focus, provides room to inspect intermediate outputs, and significantly reduces the likelihood that LLMs recognize the task as a direct search for IVs.

Prompt Construction Techniques

The effectiveness of the method relies heavily on specific prompt engineering strategies:

We propose using multi-step and role-playing prompting strategies are effective for simulating the endogenous decision-making processes of economic agents.

Role-playing prompts are proposed because they align with the source of endogeneity, which is often an agent's decision. The authors detail specific prompt templates for different tasks, including:

"Prompt 1 (Search for IVs): you are [agent] who needs to make a [treatment] decision in [scenario]. what are factors that can determine your decision but do not directly affect your [outcome], except through [treatment] (that is, factors that affect your [outcome] only through [treatment])? list [K 0] factors that are quantifiable. explain the answers."

"Prompt 2 (Refine IVs): you are [agent] in [scenario], as previously described. among the [K 0] factors listed above, choose [K] factors that are most likely to be unassociated with [confounders], which determine your [outcome]. the chosen factors can still influence your [treatment]. for each chosen factor, explain your reasoning."

Application and Extensions

The method was applied to three well-known examples in economics: returns to schooling, supply and demand, and peer effects. The results demonstrated that GPT-4 produced lists of candidate IVs that included both variables already popular in the literature and potentially new ones with rationale for their validity. Furthermore, the strategy was extended to finding (i) control variables in regression and difference-in-differences methods and (ii) running variables in regression discontinuity designs. The authors also introduce an adversarial LLM approach to review and refine discovery processes, mimicking human researcher critique.

Conclusion

The paper concludes that while the discovered lists are not absolute, they serve as a valuable benchmark that inspires empirical researchers. The overall proposal is to systematically “search for exogeneity,” positioning AI not just as a data-processing tool but as a means to improve conventional research practices by assisting in the creative and heuristic parts of causal inference. Future work could involve exploring multiple LLM agents or fine-tuning open-source models.

IVs Suggested Rationale Provided Citations

(The paper provides extensive tables detailing the suggested IVs and their rationales for various economic examples, such as returns to college attendance, production function estimation, and demand estimation.)

(Table 1 through Table 8 present specific candidates for IVs across these applications.)

Key References Cited:

(The paper references foundational work in causal inference by Imbens and Angrist (1994), Pearl (2000), Heckman (1979), and others, alongside recent work on LLMs in causal discovery.)

(References include: Card (1995, 1999, 2001); Angrist & Krueger (1984); Berry et al. (1995); Chernozhukov et al. (2024).)

Improvements for AI systems

Here are specific improvements to AI systems based on the principles and methods outlined in the paper, along with what an improved system could achieve:


  1. Acknowledge and Systematize Heuristic Search: The AI should move beyond simple retrieval or hypothesis generation by explicitly modeling the heuristic nature of human discovery.

  2. Implement Multi-Step, Role-Playing Prompting: The system should be architected to execute a structured, multi-stage prompting workflow (e.g., Step 1 for Relevance/Exclusion, Step 2 for Independence) using role-playing prompts that simulate agent decision-making processes rather than academic discourse.

  3. Integrate Counterfactual Reasoning as a Core Mechanism: The AI should be trained to explore alternative scenarios and counterfactual statements explicitly (as seen in Prompt 2), allowing it to test the Exclusion assumption (Assumption EX) by imagining outcomes given specific instrumental variables.

  4. Develop an Adversarial Refinement Loop: The system should incorporate a second LLM agent acting as a critic. This adversarial loop would force the primary model to defend its IV candidates against skeptical, non-jargon-using counter-arguments, leading to more robust and sophisticated variable sets (as shown in Section 5).

  5. Enable Contextual Conditioning via Covariates: The AI should be able to seamlessly incorporate specific, realistic contextual details (covariates) into the search process by assigning specific values or defining ranges for these factors within the role-play prompts, moving beyond purely abstract theoretical searches.

  6. Systematize Search Across Causal Inference Methods: The improved system should be modular, allowing it to adapt its prompting strategy to different causal inference paradigms:

  7. Control Variables (for Regression/DiD): Prompting sequences (like C1-C2) should be integrated to systematically search for covariates that satisfy conditional independence or parallel trends assumptions.

  8. Search for Running Variables in RDDs: The system needs a dedicated module using prompts like R1 and R2, capable of searching for assignment variables with verifiable cutoffs from external sources to identify suitable running variables.

The improved AI system could achieve the following:

  1. An autonomous Causal Discovery Engine that doesn't just find existing literature but actively generates novel, context-specific instrumental variables (IVs) for complex economic problems (e.g., finding new IVs for peer effects or supply/demand).

  2. A tool capable of generating a comprehensive list of candidate control and running variables tailored to specific empirical designs like Difference-in-Differences or Regression Discontinuity, significantly reducing the manual literature review required by researchers.

  3. The ability to stress-test its own findings through an adversarial process, resulting in IVs that are not just plausible but demonstrably more robust against potential critiques from a skeptical researcher.

  4. A system that can generate highly specific prompts for LLMs, allowing researchers to guide the AI toward finding variables relevant to their exact policy questions or setting (e.g., find IVs related to climate change regulations).

  5. A unified interface where the user can switch between searching for IVs, control variables, and running variables with a consistent, structured prompting protocol that mimics expert human research workflows.

Sources

Related papers