CausalVerify: End-to-End Verification of Causal Analyses by Language Models
cs.AI, cs.CL, econ.EM
Submitted: 2026-09-07
Updated: 2026-09-26
Comments: 22 pages, 9 figures, 12 tables. Code and data: https://github.com/causalverify/causalverify
Code: https://github.com/causalverify/causalverify
License: http://creativecommons.org/licenses/by/4.0/
The gist: Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate.
Terminology
Abstract
Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretation from verifiable computation. It pairs 259 published economics papers (reconstructed research question, data description, institutional context) with 100 fixed-seed synthetic scenarios that realise CSV datasets for difference-in-differences, event study, instrumental variables, and regression discontinuity designs. Experiment A (real-paper text agreement) scores method-family and direction agreement against four-LLM consensus labels. Experiment B (synthetic execution) runs model-written R code and checks whether the extracted treatment-effect estimate matches a canonical estimator on the same realised dataset; this execution-grounded correctness layer is L2b+, distinct from L2b, which records only whether the code executes. A calibration arm asks whether self-reported confidence separates correct from incorrect workflows. On Experiment B, seven LLMs reach L2b+ pass rates of 10% to 88% at the default 50% tolerance, and 66 of the 426 workflows that execute (15.5%) return a wrong estimate. Execution ranking (L2b) agrees with L2b+ far better than text-direction scoring (L4): Kendall τ=0.81 and Spearman ρ=0.93, versus Kendall τ between-0.20 and 0.10 for L4. Llama-3.3-70B-Instruct shows the same qualitative gap, and reported confidence does not reliably separate correct from incorrect workflows. The claims are confined to standardized single-shot workflows in these four design families under the evaluated R backend and model panel; the benchmark does not measure general causal-inference ability. Code, data, cached outputs, and a datasheet are released.
Sources
- Evaluating Large Language Models Trained on Code
- PAL: Program-aided Language Models
- Measuring Coding Challenge Competence With APPS
- Can Large Language Models Infer Causation from Correlation?
- Language Models (Mostly) Know What They Know
- DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation
- Teaching Models to Express Their Uncertainty in Words
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- ReAct: Synergizing Reasoning and Acting in Language Models
- BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection