"Cause" is Mechanistic Narrative within Scientific Domains: An Ordinary Language Philosophical Critique of "Causal Machine Learning"
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "“Cause” is Mechanistic Narrative within Scientific Domains: An Ordinary Language Philosophical Critique of “Causal Machine Learning”".
Jane: The paper was written by Vyacheslav Kungurtsev, Leonardo Christov Moore, Gustav Šír and Martin Krutský from Czech Technical University in Prague and Institute of Advanced Consciousness Studies.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's going to make a lot of people in machine learning squirm a little bit. It's called “Cause” is Mechanistic Narrative within Scientific Domains: An Ordinary Language Philosophical Critique of “Causal Machine Learning.”
Jane: And Tom, I have to say, just reading that title out loud makes me want to take a deep breath. It's a mouthful, but it's also a direct challenge to a whole field that's been booming for years.
Tom: Exactly. The paper is basically saying that when we talk about "causal machine learning," we might be overpromising what the math can actually deliver. The authors come from computer science and philosophy, and they're asking a really basic question: what do we actually mean when we say one thing causes another?
Jane: And the answer, according to them, isn't a universal law. It's different in physics, different in biology, different in economics. The word "cause" does different work in each of those fields.
Tom: Right, and that's the "ordinary language" part of the title. That's a philosophical tradition that says we should look at how people actually use words in real contexts, not try to define them in the abstract.
Jane: So instead of saying "here's the one true definition of causality," they're saying "let's look at how a physicist uses the word, how a doctor uses it, how an economist uses it." And surprise, surprise, they use it differently.
Tom: And that's a problem for the machine learning folks who claim their algorithms are discovering "true" causal relationships, because the algorithm is using one specific mathematical definition and applying it everywhere.
Jane: It's like using a hammer on every problem and being surprised when the screws don't go in straight.
Tom: That's a good way to put it. The authors aren't saying the math is wrong, they're saying the claims about what the math proves are too strong.
Jane: And that's going to ruffle some feathers, because a lot of research funding and a lot of careers are built on the idea that causal machine learning is the next big thing.
Tom: Definitely. But before we get into the nitty-gritty of their argument, let's talk about who wrote this thing. The first author is Vyacheslav Kungurtsev, who's at Czech Technical University, and he's joined by researchers from the Institute of Advanced Consciousness Studies in LA.
Jane: So you've got a computer scientist and people who study consciousness and phenomenology. That's an interesting mix.
Tom: It is, and it shows in the paper. They're not just doing a technical critique, they're bringing in Wittgenstein, Quine, Kuhn, all the big names in philosophy of science.
Jane: So this isn't a paper that's going to give you a new algorithm. It's a paper that's going to make you question whether the algorithms you already have are doing what you think they're doing.
Tom: And that's exactly why we wanted to talk about it. So stick around, because next we're going to get into the actual argument they make about how causality works in different scientific fields.
Jane: And trust me, it gets interesting when they start talking about physics versus biology versus the social sciences.
Summary: Tom: So we're back, and we're still talking about “Cause” is Mechanistic Narrative within Scientific Domains. Jane, you were going to walk us through the core argument.
Jane: Right. So the paper starts by looking at how causal machine learning actually works. It's based on these things called Bayesian networks, which are basically graphs where each node is a variable and each arrow means "this variable influences that one."
Tom: And the algorithm tries to learn the structure of that graph from data, using statistical tests for conditional independence.
Jane: Exactly. And the key assumption is something called "causal sufficiency," which means you've measured all the common causes of your variables. Nothing hidden is messing things up.
Tom: And that assumption, as the paper points out, almost never holds in the real world. Especially in biology or social science, there are always hidden confounders.
Jane: Right. But the bigger issue is more philosophical. The paper argues that the very idea of "cause" is tied to the specific scientific domain you're working in. In physics, cause is embedded in differential equations. The equations don't just describe what happens, they define the mechanism.
Tom: So when a physicist says "gravity caused the apple to fall," that's shorthand for a whole system of equations that fully describes the situation.
Jane: And in that context, causal machine learning can actually work, because the math matches the physics. The graph structure can correspond to the actual forces at play.
Tom: But then you move to biology, and things get messy. The paper calls biological systems "open and irreducible." You've got multiple scales, from molecules to cells to organs to whole organisms, and causes at one scale don't cleanly map to causes at another.
Jane: And in the social sciences, it's even worse. Human behavior is so complex that no set of equations can capture it. So what do researchers do? They tell stories. They create narratives that make sense of the data.
Tom: And that's where the title comes in. The paper says "cause" is really a mechanistic narrative within a scientific domain. It's a story that connects inputs to outputs in a way that makes sense to the experts in that field.
Jane: So the smoking and lung cancer example is perfect here. It wasn't one study that proved it. It was decades of research across epidemiology, biochemistry, physiology, all pointing in the same direction.
Tom: And each field contributed its own piece of the causal story. The epidemiologists showed the statistical association, the biochemists showed the mechanism, the pathologists showed the tissue damage.
Jane: And only when all those pieces fit together did the scientific community say "yes, smoking causes cancer." That's the preponderance of evidence approach.
Tom: So the paper's critique is that causal machine learning tries to skip all that. It tries to jump straight from data to causal claims without the mechanistic understanding.
Jane: And that's why the authors say we should be much more careful about the language we use. Instead of saying "this algorithm discovered that A causes B," we should say "this algorithm found a statistical pattern consistent with A influencing B."
Tom: It's a humility thing. And it's going to be controversial, because a lot of people in the field really do believe they're finding true causes.
Jane: They might be, sometimes. But the paper's point is that you can't know that from the algorithm alone. You need the domain expertise, the mechanistic understanding, the whole narrative.
Tom: So what's the fix? What are the authors actually suggesting we do differently? That's what we're going to dig into next.
Improvements: Tom: Okay, so we've established the problem. Causal machine learning overclaims. But what do the authors actually want us to do about it? That's what we're digging into now.
Jane: And I think the most practical suggestion is about language. They want researchers to stop saying "we discovered a cause" and start saying "we found evidence consistent with a causal relationship."
Tom: It sounds like a small change, but it's actually huge. Because the way we talk about results shapes how other people use them.
Jane: Exactly. And they tie this to something called "epistemic virtue," which is basically the discipline of being honest about how much you actually know.
Tom: And that's not just for the researchers. It's for the journals, the press releases, the headlines. The whole chain of communication.
Jane: But they also suggest a more positive direction. Instead of trying to replace domain expertise with algorithms, they want to see more integration. Mixed methods research that combines statistical analysis with mechanistic understanding.
Tom: So in medicine, that would mean a clinical trial plus a biochemical mechanism plus a physiological model, all working together.
Jane: Right. And they point to cognitive science as a great example of this. It's a field that brings together psychology, neuroscience, linguistics, philosophy, computer science, all studying the same phenomena from different angles.
Tom: And each angle contributes its own piece of the causal story. The neuroscientist talks about brain activation, the psychologist talks about behavior, the philosopher talks about consciousness.
Jane: And none of them has the whole picture, but together they build a richer understanding than any single field could achieve.
Tom: The paper also talks about how neuro-symbolic AI and explainable AI could help with this. Instead of just giving you a prediction, the model could give you a narrative that connects the inputs to the outputs in a way that makes sense to a domain expert.
Jane: So the AI becomes a tool for generating hypotheses, not for delivering final answers. It's a partner in the scientific process, not a replacement for it.
Tom: And that's actually a much more realistic and useful role for AI in science. Because the really hard problems, the ones that matter for human health and society, they're not going to be solved by a single algorithm.
Jane: They're going to be solved by teams of people from different fields, using different tools, and the AI can help them talk to each other.
Tom: There's also a really interesting point about clinical intuition. The paper says that practitioners who work with a system every day often have knowledge that isn't captured in any formal model.
Jane: And that knowledge should be respected, not dismissed because it doesn't fit into a statistical framework.
Tom: So the improvements here are about humility, integration, and respect for different kinds of knowledge. It's a much more collaborative vision of science.
Jane: And it's a vision where AI is a powerful tool, but not an oracle. That's a much healthier relationship.
Tom: So let's wrap this up. We've talked about the critique and the suggestions. What's the big picture here? That's coming up in our final segment.
Conclusion: Tom: Alright, we're at the end of our discussion of “Cause” is Mechanistic Narrative within Scientific Domains: An Ordinary Language Philosophical Critique of “Causal Machine Learning.” Let's pull it all together.
Jane: So the paper's core message is that causality isn't a single thing you can discover with one algorithm. It's a concept that takes on different meanings in different scientific contexts.
Tom: And the danger is when we take a tool designed for one context and apply it to another without thinking about whether the assumptions hold.
Jane: The authors aren't saying causal machine learning is useless. They're saying it's being oversold. The methods are powerful, but they don't deliver what they promise.
Tom: And the fix is to be more careful with our language, more humble about our claims, and more willing to combine statistical tools with deep domain expertise.
Jane: It's a call for a more mature, more integrated approach to science. One where AI is a collaborator, not a replacement for human understanding.
Tom: And honestly, that's a message that applies beyond just causal machine learning. It's about how we think about AI in general.
Jane: Yeah, because the temptation to overclaim is everywhere. It's not just in causal inference, it's in all of AI research.
Tom: So we should be grateful to the authors for reminding us to slow down and think about what we're actually saying when we make claims about what our models can do.
Jane: And to remember that the word "cause" carries a lot of weight. It's not just a statistical relationship, it's a statement about how the world works.
Tom: And that statement deserves to be backed up by more than just a p-value.
Jane: Exactly. So that's our take on this paper. It's a challenging read, but a valuable one for anyone working in machine learning, statistics, or any field that uses those tools.
Tom: And with that, we're going to say goodbye to this paper and get ready for the next one. Thanks for listening, everyone.
Jane: See you next time.
Vyacheslav Kungurtsev, Leonardo Christov Moore, Gustav Šír, Martin Krutský
Czech Technical University in Prague · Institute of Advanced Consciousness Studies
cs.LG
Submitted: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 42/100
The gist: The paper critiques the premise of causal learning in statistics and machine learning by examining the epistemology of causality across scientific disciplines, applying the Ordinary Language method
Key concepts
- Causal Machine Learning
- A field that attempts to discover 'true' causal relationships using algorithms like Bayesian networks. The paper critiques this approach, arguing that the math often overpromises what can be delivered regarding causality.
- Ordinary Language Philosophy
- A philosophical tradition that suggests understanding a word's meaning by observing how people actually use it in real-world contexts, rather than trying to define it abstractly or universally.
- Mechanistic Narrative
- The paper argues that 'cause' is fundamentally a narrative—a story connecting inputs to outputs—that makes sense within a specific scientific domain (like biology or medicine).
- Domain Expertise
- Deep knowledge of a specific field (e.g., biochemistry, epidemiology) that is necessary to interpret data and understand the context of relationships. The paper stresses this expertise cannot be replaced by algorithms alone.
Terminology
Summary
The paper critiques the premise of causal learning in statistics and machine learning by examining the epistemology of causality across scientific disciplines, applying the Ordinary Language method of anthropological investigation of customary word use to investigate valid semantics of reasoning about cause and effect in the real world.
The authors observe that although cause-and-effect semantics vary between scientific domains, they maintain a consistent central function of describing the mechanisms underlying forces most salient to the systems the scientific domain studies.
They demarcate three categories of scientific domains based on the degree to which mathematical representations of models exhibit faithful correspondence to expert knowledge in mechanism: "1) physics and engineering as domains wherein mathematical models are sufficient to comprehensively describe causality, in contrast to 2) biology, which studies open and irreducible systems with mechanisms crossing scales through emergence, and 3) the social sciences as suffering from compounding difficulties for precision but providing, through Hermeneutics, the potential for subjective phenomenology yielding findings instrumentally useful when interpreted through individual thought schemas and compelling narratives."
The authors posit that "the greater the discrepancy between expert-defined models and complete characterization of mechanism, the more that epistemic virtue requires that definitive causal claims regarding phenomena can only come through an agglomeration of consistent evidence across multiple domains."
The paper examines the formal foundations of causal inference and discovery, including Bayesian Networks, Dynamic Bayesian Networks, Structural Causal Models, Additive Noise Models, and CP-logic. The authors note that "contemporary causal inference has known fundamental issues. First, in practice, it is very rare for a paper performing causal learning on real data to identify a graph with all conditional independence hypothesis tests passing... Next, the fundamental assumption of causal sufficiency does not hold for most phenomena of interest."
The authors argue: "The fundamental matter of concern as far as the use of terminology to the effect of 'Causal Learning' and 'Causal Inference' is not that there is no validity at all to the statistical methods involved as far as being relevant to the understanding of causes. Rather, the bewitchment comes because the computation of a particular DAG and set of conditional independence hypothesis tests do not conclusively present evidence for a causal relationship between variables... And, fundamentally, nor can it. Since 'cause' is anthropologically contingent, as far as a scientific language game, a causal inference can only be made in the logic of that scientific discipline, and establishment of causes in the accepted practice of establishing positive new results on the mechanisms in the field."
The authors state: "The phrase 'this set of numerical results imply that A is a cause for B' is not correct, rather, they may suggest evidence towards that possibility. As it is, Causal Learning is in the language game of learning techniques, which is computational mathematics. They suggest that causal learning is more accurately described as
a Statistical Model with Interpretable Representations of Object Relationships."
The paper reviews historical philosophical perspectives on causality, including Descartes' a priori arguments, Hume's skepticism about induction, Kant's response positing causality as a priori knowledge, Popper's falsifiability, Kuhn's paradigms, Lakatos's research programs, and Quine's critique of the analytic-synthetic distinction. The authors note that one cannot perform induction towards causality without significant inductive bias, corresponding in practice to model representation in technical induction.
The authors state: "In Physics there are mathematical foundations that completely and comprehensively express theoretical physics' model of the world. That is, while narrative explanations of physical intuition are helpful for understanding, all assertions regarding cause and effect in the physics domain can be expressed in a formalism of equations without any essential component missing. They note that
causation can be understood as sensitivity perturbation results on the equations corresponding to the system. Engineering is described as
applied physics and
a clinical branch of Physics."
The paper describes how biological systems do not satisfy either criteria
of closed systems with well-developed theoretical models. A human body is influenced by an innumerable set of factors throughout its lifetime, making a completely identified model representation impossible, as far as including all possible causal influences.
The authors describe the standard research protocol in medical science as an iterative process: initial observational findings, followed by domain expert consideration of potential mechanisms, targeted observational studies, randomized clinical trials, in vitro experiments, and biophysical models. They present smoking and health outcomes as a case study where the collection of empirical and domain science evidence accumulated
across multiple disciplines to establish causation.
The authors argue that the layer of emergence present in biology cascades, in the fields of psychology, sociology, economics, etc. into monumental intractability.
They note that "recent events have discredited the social science establishment and its authority in the domain - the omnipresence of a failure of replication and reproducibility has affected even hard core paradigmatic textbook theses of these fields."
They present two central theses: "1) On the one hand, much greater humility should be practiced on the part of social scientists as far as the strength of certainty expressed with any claim of general mechanisms of cause and effect... 2) On the other hand, while the grand collection of established predictive claims in the social sciences must be reduced by an order of magnitude, this does not reduce social sciences' utility to a commensurate degree... micro utility can and should be maximized – social science can provide guides and expansive catalogues of models, narratives, pictures, critiques, and guidance for appropriating narratives in literature, religion, mythology, etc. to assist individuals find their truth as far as living a life that is full of joy, meaning, fulfillment, and overall flourishing in the mental space."
The paper presents cognitive science as a crystallizing contemporary example of a comprehensively multi-disciplinary field.
It notes that "Cognitive science is a unique scientific field in that, with relatively equal emphasis, it studies both the substrate and functions of the mind from the exterior, as well as the experience of that substrate - itself having a mind. The authors argue that the
hard problem of consciousness
is actually an opportunity for a sizable number of fields... to be put to work to each contribute to the understanding of the qualia to brain tissue activity correspondence."
The authors conclude: "We hope that this work provides sufficient force of argument for exhibiting greater intellectual humility in the sciences. At the same time, we hope that rather than engendering disappointment and demotivating scientific enterprise, it recognizes that the challenge of epistemic virtue in expressing degrees of certainty provides a rich Research Program in and of itself. Practice becomes less about debate in conflict, but harmonization and integration."
They note that "exercising greater caution in communicating the degree of certainty evidence provides, especially demanding restraint in the face of incentives to overstate research conclusions, is the only durable solution to modern science's collective action problems. The paper also suggests that
these investigations into multidomain causality languages present new approaches for explainability in AI, and that
Using techniques in neuro-symbolic modeling, one can use logics to define and integrate knowledge extracted from narrative descriptions and highly parameterized empirical models associated with particular domains."
Improvements for AI systems
Based on the paper's central thesis—that cause
is a mechanistic narrative within specific scientific domains, not a universal statistical property—I can implement the following concrete improvements to AI systems:
Improvement: Add a causal confidence gating
layer to any AI system that outputs causal statements (e.g., medical diagnosis AI, policy recommendation engines, scientific discovery tools).
Implementation: Before the system says X causes Y,
it must classify the domain (physics, biology, social science) and then:
-
For physics/engineering: allow direct causal language only if the model is a closed-form differential equation system with validated parameters.
-
For biology/medicine: require the system to output
consistent with a causal mechanism
and list at least two independent scales of evidence (e.g., molecular pathway + population-level trial). -
For social science: force the system to output
suggests an association
and explicitly prohibit the wordcause
unless the user provides a multi-domain narrative (e.g., historical + psychological + economic).
What the improved AI can do: It will no longer overstate findings from observational data. For example, a healthcare AI will say This drug is associated with reduced mortality in this cohort, but a causal claim requires confirmation via a randomized trial and a biochemical mechanism
instead of This drug reduces mortality.
Improvement: Build an AI module that, given a causal hypothesis, automatically searches and integrates evidence from at least three different language games
(e.g., statistical, mechanistic, and narrative/hermeneutic).
Implementation: Use a retrieval-augmented generation (RAG) system that:
-
Queries statistical databases (e.g., clinical trials, econometric studies)
-
Queries mechanistic literature (e.g., biochemistry, neuroscience, physics equations)
-
Queries qualitative/narrative sources (e.g., ethnographic studies, historical accounts, patient testimonials)
-
Then outputs a
preponderance of evidence
score, not a single p-value.
Improvement: For AI systems used in psychology, sociology, or policy, add a narrative coherence
check that evaluates whether a causal explanation is compelling and instrumentally useful to a human, not just statistically significant.
Implementation: Use a fine-tuned language model trained on therapeutic narratives (e.g., CBT, psychodynamic, ACT) and historical/sociological treatises (e.g., Weber, Bourdieu). The AI will:
-
Generate multiple causal narratives for the same phenomenon (e.g.,
depression is caused by cognitive distortions
vs.depression is caused by unresolved childhood conflict
) -
Score each narrative on internal consistency, alignment with known facts, and potential for positive behavioral change in the user.
-
Present the user with the narrative that is most actionable for their specific context, rather than claiming one is objectively
true.
Improvement: For AI systems in engineering or physics (e.g., autonomous vehicle control, climate modeling), enforce that any causal claim must be derived from a known physical law (ODE/PDE) and not from a purely data-driven graph.
Implementation: Modify the training loss of any neural network that outputs causal relationships to include a physics-based penalty term. For example, if the network suggests A causes B,
the system must verify that there exists a plausible differential equation linking A and B (e.g., F=ma, Maxwell's equations). If not, the output is rejected or downgraded to correlational.
What the improved AI can do: A self-driving car's decision system will only attribute the pedestrian caused the braking
if the kinematic equations (distance, velocity, reaction time) are satisfied. It will never say the red car caused the accident
based solely on sensor correlation without a physical mechanism.
Improvement: Add a certainty calibration
module to AI research assistants (e.g., tools that summarize literature or propose hypotheses).
Implementation: The AI will output a causal certainty score
from 0–100, where:
-
0–30:
We observe a correlation; no mechanism known.
-
31–60:
A plausible mechanism exists, but only at one scale (e.g., statistical or mechanistic, not both).
-
61–85:
Evidence converges across at least two independent domains.
-
86–100:
Closed-form physical law or replicated randomized trial with a known biological mechanism.
Improvement: Implement a neuro-symbolic module that uses CP-Logic to transfer causal knowledge across domains via analogy, as suggested in the paper's conclusion.
Implementation: The system will:
-
Represent causal rules in CP-Logic (e.g.,
Cold causes shivering in living organisms
) -
Use a structure-mapping engine to find analogies (e.g.,
Slow molecular motion corresponds to cold; kinetic energy input corresponds to shivering
) -
Generate new hypotheses for the target domain (e.g.,
Injecting kinetic energy into slow-moving molecules may cause them to become more active
) -
Then validate the hypothesis using the multi-scale evidence engine (improvement #2).
-
Avoid overclaiming: It will never say
X causes Y
without domain-appropriate evidence. -
Integrate evidence: It will automatically combine statistical, mechanistic, and narrative evidence.
-
Respect domain semantics: It will use
cause
only in physics/engineering contexts with closed-form equations; elsewhere it will useassociated with
orconsistent with a mechanism.
-
Support human decision-making: In social sciences, it will provide multiple compelling narratives and let the user choose, rather than forcing a single
truth.
-
Transfer knowledge safely: It can propose analogical causal hypotheses and then test them rigorously across scales.
These changes directly address the paper's core critique: that causal machine learning overstates its certainty by ignoring the domain-specific nature of cause.
The improved AI is more honest, more useful, and less likely to cause harm through misplaced confidence.
Abstract
Causal Learning has emerged as a major theme of research in statistics and machine learning in recent years, promising computational techniques to reveal ``true'' causality. In this paper, we critique the premise of causal learning by considering the epistemology of causality across disciplines, applying the Ordinary Language method of an anthropological investigation of customary word use in reasoning about cause and effect in the real world. We observe that although cause-and-effect semantics vary between scientific domains, they maintain a consistent central function of describing the mechanisms underlying forces most salient to the systems under research. A critical distinction is the degree to which the mathematical representations of the models used for statistical analysis exhibit a faithful correspondence to expert knowledge in mechanism. We demarcate 1) physics and engineering as domains wherein mathematical models are sufficient to comprehensively describe causality, in contrast to 2) biology, which studies open and irreducible systems with mechanisms crossing scales through emergence, and 3) the social sciences as suffering from compounding difficulties for precision but providing, through Hermeneutics, the potential for subjective phenomenology yielding findings instrumentally useful to individuals. We posit the greater the discrepancy between expert-defined models and complete characterization of mechanism, the more that epistemic virtue requires that definitive causal claims regarding phenomena can only come through an agglomeration of consistent evidence across multiple domains. Exercising greater caution in communicating the degree of certainty evidence provides, especially demanding restraint in the face of incentives to overstate research conclusions, is the only durable solution to modern science's collective action problems.
Sources
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Rule Extraction Algorithm for Deep Neural Networks: A Review
- Opening the Black Box of Deep Neural Networks via Information
- Respecting causality is all you need for training physics-informed neural networks
- Marrying Causal Representation Learning with Dynamical Systems for Science
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks