Causal Agent based on Large Language Model

arXiv:2408.06849 · cs.AI, cs.CL · Submitted 2026-08-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Causal Agent based on Large Language Model".

Jane: The paper was written by Kairong Han, Kun Kuang, Ziyu Zhao, Junjian Ye and Fei Wu from Zhejiang University and Huawei Technologies Company.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we’re looking at a paper that’s been making waves on arXiv, and it’s called “Causal Agent based on Large Language Model.” Jane, I have to say, the title alone got me curious — we’re talking about giving LLMs the ability to actually reason about cause and effect, not just pattern-match.

Jane: Absolutely, Tom. And the authors — Kairong Han, Kun Kuang, Ziyu Zhao, Junjian Ye, and Fei Wu — they’re from Zhejiang University and Huawei. That’s a solid mix of academia and industry. The core idea is that LLMs are great at language, but when you hand them a spreadsheet of data and ask “does smoking cause lung cancer?” they kind of freeze.

Tom: Right, and that’s the gap they’re trying to close. The paper argues that causal reasoning is inherently hard for LLMs because causal problems are usually about tabular data, not natural language. And the theory itself is hard to describe in words. So they built an agent that uses tools to bridge that gap.

Jane: And that’s where the “agent” part comes in. Instead of asking the LLM to reason directly, they give it access to causal analysis libraries — like causal-learn and EconML — and let it call those tools to analyze the data. The LLM becomes the brain that decides which tool to use and how to interpret the results.

Tom: So it’s like giving a chef a kitchen full of appliances instead of asking them to cook with their bare hands. The LLM still does the thinking, but it has the right equipment to actually get the job done.

Jane: Exactly. And the implications are huge. If this works, we could automate a lot of scientific discovery — not just in medicine, but in economics, social science, anywhere you need to figure out what actually causes what. That’s a big deal.

Tom: And the authors are careful to model the problem at four levels — variable, edge, graph, and effect — so it’s not just one trick. They’re building a whole framework for causal reasoning. I’m excited to see how they tested it.

Jane: Me too. Let’s get into the details of how they actually structured the problem and what the results look like. That’s coming up next.

Summary: Tom: So we’re back with “Causal Agent based on Large Language Model,” and Jane, we were just talking about the four levels of causal problems they defined. Can you break those down for our listeners?

Jane: Sure, Tom. The first level is the variable level — that’s basically asking whether two variables are correlated or independent. The second is the edge level, where you ask about direct causal relationships, like whether A directly causes B. Then there’s the causal graph level, where the agent has to generate a whole network of cause-and-effect relationships. And finally, the causal effect level, where you’re quantifying things like the average treatment effect — how much does changing one variable change another.

Tom: And they built a benchmark called CausalTQA to test all of this. About one thousand four hundred questions across those four levels. What did they find?

Jane: The results are pretty impressive. At the variable level, they’re hitting over ninety-five percent accuracy on all three sub-problems. At the edge level, it’s over eighty-nine percent. The causal graph level is a bit lower — around eighty-one percent for full graphs and ninety-one percent for partial graphs — but still strong. And the causal effect level is at ninety-eight percent.

Tom: That’s really strong. But here’s what I find interesting — they compared against a baseline called code-LLM, where the model just writes Python code to analyze the data. And that baseline performs terribly, sometimes even worse than random.

Jane: Right, and that’s the key insight. It’s not that LLMs can’t reason about causality at all — it’s that they need the right scaffolding. When you give them tools and a structured way to use them, they can do really well. The code-LLM baseline struggles because it has to figure out everything from scratch, and it makes all sorts of errors — wrong variable names, syntax errors, logic mistakes.

Tom: So the agent framework is doing the heavy lifting. The LLM is still the decision-maker, but it’s working with reliable tools instead of trying to reinvent the wheel every time.

Jane: Exactly. And they also tested on a real-world dataset called QRData, and their agent beat the previous state-of-the-art by six percent. That’s a big deal because it shows the approach generalizes beyond synthetic data.

Tom: And that’s the part that gets me excited — this isn’t just a lab experiment. It’s actually working on real problems. But I’m curious about how the agent actually decides which tools to use and in what order. That’s the planning part, right?

Jane: Yeah, that’s the ReAct framework — the agent thinks, acts, observes, and repeats. It’s like a scientist running experiments: you try something, see what happens, and adjust. We should talk about how that planning process works in more detail.

Improvements: Tom: We’re back with “Causal Agent based on Large Language Model,” and Jane, we’ve covered the results, but I want to dig into the actual improvements this paper proposes. What makes this agent different from just giving an LLM a Python interpreter?

Jane: Great question, Tom. The biggest improvement is the memory module. Instead of storing everything as text, the agent keeps a dictionary where the keys are names and the values are actual causal graph objects. So when the agent generates a causal graph, it can save it and reference it later without having to regenerate it.

Tom: That’s clever. So it’s like having a whiteboard where you can write down intermediate results instead of trying to remember everything in your head.

Jane: Exactly. And that’s crucial because solving a complex causal problem might require multiple steps — first generate the graph, then check if there’s a confounder, then maybe estimate a causal effect. If you lose the graph between steps, you have to start over.

Tom: And that’s where the planning module comes in. The agent uses a ReAct-style loop — it thinks about what to do, takes an action, observes the result, and decides whether it can answer the question or needs to keep going. On average, it takes about two point five iterations per question.

Jane: And the tools themselves are re-encapsulated so they accept JSON input and produce natural language output. That’s a subtle but important improvement — the LLM doesn’t have to parse complex data structures. It just gets a clear, readable result.

Tom: So it’s like the difference between giving someone a manual in a foreign language versus translating it into their native tongue. The tools speak the LLM’s language.

Jane: Right. And they also use in-context learning with one-shot examples to show the agent how to use the tools. That helps it understand the format and the intent of each tool before it starts.

Tom: I also noticed they tested with different LLMs as the decision core — GPT-three point five, ChatGLM-four-air, and ChatGLM-four-plus. The stronger the model, the better the results. So the framework scales with the underlying model’s capabilities.

Jane: That’s a good sign. It means the framework isn’t tied to one specific model — it’s a general approach that can improve as LLMs get better. And that’s the kind of forward-looking design you want to see.

Tom: And they even compared against RestGPT, another agent framework, and their causal agent did much better. That suggests the specialized design matters — you can’t just slap any agent framework on top of causal tools and expect it to work.

Jane: Exactly. The causal graph as an intermediary — that’s the secret sauce. It ties everything together and lets the tools share information in a meaningful way.

Tom: So what does this mean for the future? I’m thinking about real-world applications — healthcare, economics, policy-making. Where do you see this going?

Jane: I think the potential is enormous. Imagine an AI assistant that can analyze clinical trial data and tell doctors which treatment actually works, or a tool that helps economists understand the real drivers of inflation. That’s the kind of impact this could have.

Conclusion: Tom: We’ve had a great conversation about “Causal Agent based on Large Language Model,” and I think we’ve only scratched the surface. Jane, what’s the big takeaway for our listeners?

Jane: The big takeaway is that LLMs don’t have to be bad at causal reasoning. They just need the right tools and the right framework. By giving them access to causal analysis libraries and structuring the reasoning process, we can turn them into powerful causal reasoners — and that’s a huge step forward.

Tom: And it’s not just about the results, which are impressive — over ninety-five percent accuracy on variable-level questions, ninety-eight percent on causal effect estimation. It’s about the approach. The agent is transparent, interpretable, and reliable. You can see every step it takes, which is crucial when you’re making decisions that affect people’s lives.

Jane: Absolutely. And the fact that it generalizes to real-world data — beating the previous state-of-the-art by six percent on QRData — shows that this isn’t just a synthetic benchmark trick. It works in practice.

Tom: I also love that the authors made their code available on GitHub. That’s going to accelerate research in this area. Other teams can build on this framework and push it even further.

Jane: And that’s the exciting part — this is just the beginning. As LLMs get better and causal tools get more sophisticated, we’re going to see agents that can tackle even more complex causal problems. Maybe even ones that can design their own experiments or discover new causal relationships we hadn’t thought of.

Tom: Well said, Jane. So let’s say goodbye to “Causal Agent based on Large Language Model” and get ready for the next paper. This one’s going to be a hard act to follow.

Jane: Agreed. Thanks for listening, everyone. We’ll be back soon with more exciting research from arXiv.

Tom: Until next time, keep asking why — and maybe someday, your AI assistant will be able to answer.

Kairong Han, Kun Kuang, Ziyu Zhao, Junjian Ye, Fei Wu

Zhejiang University · Huawei Technologies Company

cs.AI, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/kairong-han/causal_agent

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 65/100

The gist: The paper "Causal Agent based on Large Language Model" addresses the challenge of enhancing large language models' (LLMs) causal reasoning capabilities.

Key concepts

Causal Agent
This is a framework where an LLM acts as the decision-maker for solving causal problems. Instead of reasoning directly, it uses tools like 'causal-learn' to analyze data and interpret results, allowing it to perform complex analysis.
ReAct Framework
The agent uses this planning process: Think about a step, take an action (using a tool), observe the result, and repeat. This iterative loop allows the LLM to solve problems by testing hypotheses and adjusting its approach.
Causal Levels
The paper models causal problems at four levels: variable level (correlation), edge level (direct cause-and-effect), graph level (generating a whole network of causes), and effect level (quantifying how much one variable changes another).
Code-LLM Baseline
This is a standard test where an LLM simply writes Python code to analyze data. The study found this baseline performs poorly, highlighting that the LLM needs structured tools and a framework to succeed.

Terminology

Summary

The paper Causal Agent based on Large Language Model addresses the challenge of enhancing large language models' (LLMs) causal reasoning capabilities. The authors identify two fundamental limitations: the inherent complexity of causal problems and causal theory poses challenges in accurately describing them in natural language, making it difficult for LLM to comprehend and use them effectively, and causal datasets are typically tabular data, while large models excel in handling natural language data... This structural heterogeneity hinders LLM from effectively reasoning with tabular data.

To address these challenges, the authors propose a causal problem modeling approach from the perspective of the LLM and propose a causal agent framework by guiding LLM to invoke causal tools. They model causal problems into four levels: variable level, edge level, causal graph level, and causal effect level. The variable level focuses on the agent's judgment and understanding of correlations, the edge level focuses on the agent's examination of causal relationships between variables, the causal graph level focuses on the agent's ability to generate causal graphs, and the causal effect level focuses on the agent's estimation of causal effects between variables for quantitative expression.

The causal agent framework consists of three modules: tools, memory, and plan. In the tools module, the causal agent invokes the causal analysis library in Python programming tools, such as causal-learn and EconML, enabling it to receive a pair of tabular data and a causal problem description of the data as input. In the plan module, the causal agent utilizes its text comprehension and reasoning abilities to obtain answers to causal problems in many iterations, inspired by the ReAct framework. In the memory module, the agent maintains an instantiated dictionary where the keys are names and the values are causal graphs, which allows the agent to retrieve the necessary causal graph using the key.

To verify the causal agent's capabilities, the authors established "a Causal Tabular Question Answer (CausalTQA) benchmark consisting of four levels of causal problems: variable level, edge level, causal graph level, and causal effect level. CausalTQA consists of about 1.4K for these four levels questions. The benchmark was generated using non-linear additive Gaussian noise models for tabular data, with questions generated from 81 question templates and data is sampled from a variable pool consisting of 311 variables."

Experimental results show that "the causal agent demonstrates a high accuracy in correctly invoking tools and producing the expected answers at four level questions, with accuracy rates of over 95% in all three subproblems for determining correlation at the variable level, over 89% in all three sub-problems at the edge level, over 81% in the causal graph level, and 98% in the causal effect estimation level." Specifically, the variable level achieved 95.1% accuracy on independent tests, 99.4% on conditional independence tests, and 99.4% on multi-conditional independence tests. The edge level achieved 89.5% accuracy for direct cause relationships, 97.4% for colliders, and 94.6% for confounders. The causal graph level achieved 81.8% accuracy for complete graphs and 91.6% for partial graphs.

The authors also compared their approach against a baseline (code-LLM) where we let code-LLM write code for causal problems and run the code once in the programming environment. The causal agent significantly outperformed this baseline across all levels. Additionally, Through verification on the real-world dataset QRData, the causal agent is 6% higher than the original SOTA, demonstrating its strong generalization ability in real-world scenarios, with the causal agent achieving 57.3% accuracy using GLM-4-Plus compared to the previous SOTA of 51.3%.

The paper's contributions include: A hierarchical modeling perspective has been proposed for LLM to solve causal problems, The causal agent has been proposed to empower LLM with the ability to solve causal problems, high accuracy results on the CausalTQA benchmark, and demonstration of the potential of the causal agent for automated causal reasoning on tabular data.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system:

  • Add a tool-calling interface that connects the LLM to causal analysis libraries (causal-learn, EconML)

  • Implement JSON-formatted tool inputs/outputs to bridge tabular data with natural language

  • Create encapsulated functions for: independence testing, conditional independence testing, PC algorithm for causal graph generation, and LinearDML for causal effect estimation

  • Implement a 4-level causal problem solver:

  • Variable level: Correlation/independence detection (direct, single-conditional, multi-conditional)

  • Edge level: Causal relationship classification (direct cause, confounder, collider)

  • Graph level: Full and partial causal graph generation

  • Effect level: Average treatment effect (ATE) estimation

  • Maintain a dictionary-based memory where keys are causal graph names and values are Python class instances of causal graphs

  • Enable cross-tool information sharing through graph references rather than text descriptions

  • Allow retrieval of intermediate causal graphs for multi-step reasoning

  • Implement multi-round thought-action-observation loops where the system:

  • Generates a thought about what causal analysis to perform

  • Executes the appropriate causal tool

  • Observes results and determines if the original question is answered

  • Iterates up to 15 times maximum

  • Add prompt templates that handle both medical and market domain causal questions

  • Include in-context learning examples showing proper tool usage

  • Implement answer constraint (yes/no/uncertain) for variable and edge level questions

  1. Answer causal questions about tabular data with 95%+ accuracy on correlation tests, 89%+ on causal edge detection, 81%+ on causal graph generation, and 98%+ on causal effect estimation

  2. Automatically determine:

  • Whether variables are independent or conditionally independent

  • Direct causal relationships, confounding factors, and collider structures

  • Complete or partial causal graphs from data

  • Average treatment effects with specified covariates

  1. Handle real-world data with 6% higher accuracy than previous state-of-the-art on the QRData benchmark

  2. Provide transparent reasoning by showing intermediate causal graphs and tool outputs, making the decision process auditable

  3. Scale with model capability – performance improves when using stronger base LLMs (e.g., GLM-4-Plus achieves near-perfect accuracy on most tasks)

  4. Process tabular data without manual preprocessing – the system handles the alignment between natural language questions and structured data automatically through tool invocation

Sources

Related papers