Position: Reasoning is a Learnable Rule-Based Process

arXiv:2608.12325 · cs.AI, cs.CL, cs.LG · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Position: Reasoning is a Learnable Rule-Based Process".

Jane: The paper was written by Rachel Lawrence and Jacqueline Maasch from Microsoft Research and Cornell Tech.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a title that just grabs you: "Position: Reasoning is a Learnable Rule-Based Process." Jane, I have to say, just reading that title got me thinking—are we about to see a big fight between the old-school symbolic folks and the modern deep learning crowd?

Jane: Oh, absolutely, Tom. And that's exactly why I love this one. The paper is basically saying, "Hey, we've been arguing about what reasoning even *is* for years, and that's a problem." They're not trying to pick a side in the symbolic-versus-neural war. Instead, they're saying we need a clear, operational definition so we can actually measure progress. It's a position paper, so it's less about new experiments and more about setting the rules of the game.

Tom: Right, and they lay out three core theses right up front. The first one is basically "define, then measure." You can't just say a model is good at reasoning if you haven't told me what you mean by reasoning. It sounds obvious, but in practice, a lot of papers just throw the word around.

Jane: Exactly. And their second thesis is the big one: reasoning is a learnable rule-based process. So they're saying it's not just an output, like a correct answer on a test. It's the *process* of selecting and applying rules to your beliefs and evidence. And crucially, they say those rules can be learned, not just hard-coded by a human.

Tom: That's the bridge, isn't it? It makes the definition friendly to modern machine learning, not just to old-school logic programming. And then the third thesis is about validity—that a reasoning process is valid if it follows the rules exactly, regardless of whether the rules themselves are "good" or lead to a true answer.

Jane: Which is a really important distinction. You can have a perfectly valid reasoning process that starts from false premises and ends up with a false conclusion. That's not sound reasoning, but it's still valid in the formal sense. The paper spends a lot of time on that difference, and I think it's crucial for how we evaluate these big language models.

Tom: And that's where they introduce this fun concept of the "reasoning zombie"—a system that talks like it's reasoning but doesn't actually have the internal mechanism to back it up. It's a nod to the philosophical zombie idea, but for reasoning instead of consciousness.

Jane: I love that. It's a really concrete way to think about the problem. We can't just look at a model's output and assume it reasoned to get there. It might have memorized the answer, or guessed, or found a statistical shortcut. The paper is pushing us to look under the hood.

Tom: So, Lu, you've been quiet. What's your gut reaction to this framing? Is this the kind of definition that could actually stick?

Lu: I think it's a very brave attempt, Tom. The authors are trying to unify a lot of history—from Aristotle's syllogisms to modern Bayesian inference—under one umbrella. The fact that they can map things like reinforcement learning and even simple n-gram models into their framework is a testament to its flexibility. But that flexibility is also its biggest challenge. If everything is reasoning, does the word lose its meaning?

Tom: That's a great point to hold onto. Let's keep that in mind as we dig into the actual definition they propose.

Summary: Jane: So, we've set the stage with the title and the big theses. Now let's get into the meat of the paper, which is their formal definition. Tom, you want to take a crack at explaining it in plain English?

Tom: I'll try. They define a "state" as a snapshot of everything the reasoner knows at any given moment. That includes its current beliefs, all the evidence it has gathered, and the set of rules it has available. Reasoning, then, is the process of applying those rules to the beliefs and evidence to create a new state. And you do that over and over again until you hit a stopping rule, which is usually your goal.

Jane: And the key part is that they split the rules into two types. There are "local rules" that update your beliefs, and "meta rules" that update your other rules. So the system can literally learn new rules on the fly. That's the "learnable" part of the title. It's not a static system; it can revise its own playbook.

Lu: And that's a really elegant way to capture how modern AI works. You can think of the local rules as the model's forward pass, and the meta rules as the training algorithm that updates the weights. The paper is essentially providing a mathematical language that can describe both a simple calculator and a massive neural network.

Meng: But here's where I start to get a little skeptical. They talk a lot about "exact rule application." In a symbolic system, that's easy to verify. You run the code, you see the steps. But in a neural network, how do you even know what the "rules" are? They're buried in millions of parameters. The paper acknowledges this, but it feels like a huge gap between the theory and the practice.

Jane: That's a fair pushback, Meng. And the paper doesn't really solve that problem. But what it does do is give us a target. It says, "If you want to claim your model is reasoning, you need to be able to point to the rules and show they're being applied exactly." For a neural network, that's a really high bar, and maybe that's the point.

Tom: They even give a concrete example with a trivial Python implementation to show how the definition works in practice. It's a simple logical reasoning task, but it demonstrates how you can audit the process and check for validity. It's a proof of concept that the definition isn't just philosophical hand-waving.

Lu: And that's what I find most valuable. The paper is essentially a call to action for the community to stop treating reasoning as an undefined buzzword. They're saying, "Here's a framework. Let's argue about it, refine it, and use it to build better evaluation benchmarks." The construct validity problem is real—if we don't know what we're measuring, our benchmarks are meaningless.

Meng: So, practically speaking, this would mean that a benchmark like "GSM8K" or "ARC" would need to be redesigned to not just check the final answer, but to verify the *process* used to get there. That's a massive undertaking.

Jane: Exactly. And it connects back to their point about chain-of-thought. They argue that CoT traces are not trustworthy explanations because they might not reflect the actual internal process. A model could generate a convincing-looking reasoning trace without actually following those steps internally. So, we need better tools to look inside the black box.

Tom: And that brings us to their concept of the "reasoning zombie" again. It's not just a thought experiment. They argue that imperfect versions of these zombies already exist in the wild. Models that get the right answer on benchmarks but are actually just doing sophisticated pattern matching.

Lu: The paper suggests that we need to be very careful about when we actually *need* a true reasoner, and when a "zombie" is good enough. For a chatbot recommending a recipe, a zombie is fine. For a system making medical decisions, you absolutely need validity guarantees.

Meng: So, the next question is, what are they actually proposing we do about it? They have a checklist, right?

Improvements: Jane: Right, Meng. The paper doesn't just stop at the definition. It includes a whole checklist for best practices in AI reasoning research. It's like a code of conduct for anyone who wants to publish a paper claiming their model can reason.

Tom: And it's pretty thorough. The first section is all about defining your terms. You have to explicitly state what the rules, beliefs, and evidence are in your specific system. You can't just say "we used a transformer." You have to say, "In this setting, the rules are the learned weights, the beliefs are the hidden states, and the evidence is the input tokens."

Meng: That's a high bar, but it's the right one. It forces researchers to be honest about what their model is actually doing. The second section is about validity. You have to prove that your transitions from one state to the next are the result of exact rule application. That's where the rubber meets the road for neural networks.

Lu: And the checklist also addresses the evaluation itself. It says you need to justify the construct validity of your benchmark. You can't just say "this dataset tests reasoning." You have to explain *why* it tests reasoning according to your operational definition. This is a direct challenge to the current culture of just throwing a new dataset out there and calling it a reasoning benchmark.

Tom: There's also a big emphasis on not conflating the process with the product. Getting the right answer on a math problem is not the same as reasoning. The checklist forces you to separate those two things and evaluate them independently.

Jane: And it goes further, into the realm of communication. It says researchers should be clear about the limitations of their systems and avoid overhyping their results. That's a direct jab at the trend of press releases saying "AI achieves human-level reasoning" when the model just did well on a specific test.

Meng: So, this isn't just a theoretical exercise. This is a practical tool. If journals and conferences adopted this checklist as a requirement, it would fundamentally change how research is done. It would slow things down, but it would make the results much more trustworthy.

Lu: And that's the core of their argument. They're not saying "don't build big models." They're saying "if you're going to claim your big model reasons, you need to be able to show your work." It's about moving from a culture of "show me the accuracy" to "show me the mechanism."

Tom: They even bring up the example of Lean, the formal theorem prover. It provides a gold standard for what valid reasoning looks like because every step is formally verified. The paper isn't saying every AI system needs to be like Lean, but it's saying that if you're in a domain where correctness is critical, you should be aiming for that level of rigor.

Jane: And that's where the "learnable" part becomes so powerful. You don't have to hand-code all those rules like they did in the old expert systems. You can learn them from data, but then you need to verify that they're being applied correctly. The paper is essentially proposing a hybrid approach where you get the flexibility of deep learning and the verifiability of symbolic systems.

Meng: So, the checklist is a roadmap for building what they call "validity by design." Instead of trying to retroactively explain a black box, you build the system so that its reasoning steps are auditable from the start.

Tom: And that's a really compelling vision. It's not about going back to the 1980s. It's about taking the best of both worlds and creating a new standard for what it means to be a trustworthy AI system.

Conclusion: Jane: Well, we've had a fantastic discussion today about "Position: Reasoning is a Learnable Rule-Based Process." We started with the problem of definitional ambiguity, and we're ending with a concrete proposal for how to fix it.

Tom: That's right. The paper gives us three core ideas to take home. First, we need to define reasoning before we can measure it. Second, we should think of reasoning as a process of applying learnable rules, not just as an output. And third, validity comes from following those rules exactly, even if the rules themselves are flawed.

Lu: And the most impactful part for me is the checklist. It's a practical tool that could really change the field. It forces researchers to be explicit about their assumptions and to justify their evaluation methods. It's a push for more rigorous, more trustworthy science.

Meng: I agree. It's a high bar, but it's a necessary one. If we're going to deploy these systems in critical domains like healthcare or law, we need to know that they're not just clever parrots. We need to know they're actually following rules.

Jane: And that brings us back to the "reasoning zombie" concept. The paper warns us that we might already have these zombies among us, and we need to be able to tell them apart from true reasoners. That's not just an academic exercise; it's a matter of public safety and trust.

Tom: So, as we say goodbye to this paper, I think the biggest takeaway is a call to action. The authors are urging the community to engage with their definitions, to argue with them, to refine them. They don't claim to have the final answer, but they're providing a starting point for a much-needed conversation.

Jane: And that's what makes this paper so exciting. It's not the end of the story; it's the beginning of a new chapter. It's a challenge to all of us to be more rigorous in how we talk about and build AI reasoning.

Tom: Well said, Jane. That's all the time we have for this one. Thanks to Lu and Meng for joining us, and to all our listeners out there. We'll be back soon with another paper, but for now, this is Tom and Jane signing off.

Rachel Lawrence, Jacqueline Maasch

Microsoft Research · Cornell Tech

cs.AI, cs.CL, cs.LG

Submitted: 2026-05-29

Updated: 2026-08-14

Journal ref: Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026

Code: https://github.com/jmaasch/valid

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 63/100

The gist: The paper "Position: Reasoning is a Learnable Rule-Based Process" (arXiv:2608.12325v1, published in the Proceedings of the 43rd International Conference on Machine Learning, PMLR 306, 2026) by Rachel

Key concepts

Reasoning Zombie
A system that generates output that looks like reasoning but lacks the actual internal mechanism or process to support those steps. Hosts discuss the need to look under the hood rather than just judging the final answer.
Validity vs. Soundness
Validity means a reasoning process follows its rules exactly, regardless of whether those rules are correct. Soundness requires both following the rules (validity) and starting from true premises to reach a true conclusion.
Learnable Rule-Based Process
The paper defines reasoning not as an output, but as the process of applying rules to beliefs and evidence. Crucially, these rules can be learned from data, bridging symbolic logic with modern machine learning.
Reasoning Checklist
A proposed set of best practices for AI research. It requires researchers to explicitly define terms (rules, beliefs), prove exact rule application (validity), and justify the construct validity of their benchmarks.

Terminology

Summary

The paper Position: Reasoning is a Learnable Rule-Based Process (arXiv:2608.12325v1, published in the Proceedings of the 43rd International Conference on Machine Learning, PMLR 306, 2026) by Rachel Lawrence and Jacqueline Maasch addresses the definitional ambiguity surrounding reasoning in generative AI.

The paper identifies two main problems:

P1: "Reasoning in generative AI has experienced unnecessary and addressable definitional ambiguity, where imprecise and overloaded definitions are often misaligned with historical treatments of this topic in AI and philosophy (when definitions are provided at all)."

P2: This breeds mismeasurement, promotes an illusion of shared understanding among researchers, and subverts measurable progress toward trustworthy AI reasoning.

The paper presents three core theses:

Thesis 1: Define, then measure. Operational definitions should be stated for the reasoning phenomena under investigation, and the construct validity of reasoning evaluation should be justified with respect to these definitions.

Thesis 2: "Reasoning is a learnable rule-based process. Reasoning is a process of exact rule application, not an output. Learnable rules map reasoning inputs to outputs and can encompass theorems, functions, policies, etc., including rules pertaining to stochasticity, uncertainty, and approximation."

Thesis 3: Rule-based reasoning is valid. The validity of a reasoning process arises from exact rule application, independent of rule selection.

The paper provides an informal definition of reasoning:

Definition 2.1 (Reasoning, informal): The process of selecting and applying sequences of rules that act on prior beliefs and current evidence to obtain principled belief updates in evolving states.

Definition 2.2 (Reasoner, informal): A goal-oriented decision-maker that implements reasoning.

Definition 2.3 (State, informal): The set of all parameters that are pertinent to the reasoner at time t, including some subset of the historical record of beliefs, evidence, and rules.

The paper also provides a formal operational definition:

Definition 2.4 (Reasoning, formal): "Let St:= ⟨Bt, Et, Rt⟩ denote the reasoner's state at time step t, where Bt denotes current belief, Et denotes aggregated evidence up to time t, and Rt denotes the current set of established rules. Then, reasoning is the iterated application over steps t of rules r ∈ Rt−1 to prior beliefs Bt−1 and current evidence Et, by which we obtain dynamically updated states St, and where every output Bt for t > 0 is the result of a rule application r(Bt−1, Et) to the contents of state St−1."

The rule set is partitioned into local rules (RL, which update beliefs) and meta rules (RM, which update rules). The paper also defines validity and soundness:

Definition 2.7 (Validity): A transition from state St to St+1 is valid if and only if it arises from the application of a rule r ∈ Rt to components of state St.

Definition 2.8 (Soundness): A valid transition from state St to St+1 is sound if and only if all premises (as encoded by B, R, and E) are true with respect to external evaluation.

The paper makes several claims:

Claim 2.1: "Valid reasoning arises from exact rule application. Validity requires that each rule is always executed exactly: not partially, not approximately, not sometimes. This does not preclude rule-based means of handling stochasticity, uncertainty, and approximate inference."

Claim 2.2: "Validity is independent of rule selection. Implementing a reasoning process requires selecting which specific rule to apply at each step. Because validity is independent of soundness, and any properly-typed rule application creates a valid output, the validity of a reasoning process is independent of the algorithm used to select the rule sequence, regardless of external ground truth."

Claim 2.3: "Reasoning is commonplace. The permissiveness of Def. 2.4 may appear to undermine its value, as it admits simplistic and low-utility systems. We argue something different: when distilled to its core components, reasoning is commonplace."

Claim 2.4: A system can be simultaneously an r-zombie in one sense and a valid reasoner in another.

Claim 2.5: Rules are learnable and defeasible in the general case. We contend that rule-based reasoning and data-driven ML (e.g., probabilistic deep learning) are not mutually exclusive.

Claim 2.6: Rules are explanations. The explainability of a reasoning process lies in the rule set, as rules are the justifications by which each intermediate reasoning step is executed.

Claim 2.7: Operationalization facilitates trust. A central aspect of trust is the accurate representation of the capabilities or expected behavior of a system.

Claim 2.8: Reasoning requires memory. Notions of prior beliefs, evidence, and rules imply the existence of memory, as this body of information must be stored and recalled.

Claim 2.9: Natural language is not necessary for reasoning. Defs. 2.1 and 2.4 do not imply a necessary role of natural language in AI reasoning.

The paper introduces the concept of reasoning zombies (r-zombies): systems that superficially behave as autonomous reasoners, but lack valid internal reasoning mechanisms. The paper argues that imperfect AI r-zombies have already come into existence and that differentiating AI r-zombies from AI reasoners is theoretically and, often, empirically possible.

The paper maps its definitions to specific forms of reasoning:

Example 2.1 (Logical deduction): A natural deductive system over a formal language is initialized with a set of premises Γ, and a static set of inference rules (e.g., modus ponens or modus tollens) acting on premises.

Example 2.2 (Bayesian inference): Bayesian inference provides principled means of revising beliefs in hypotheses as new evidence emerges. We iteratively refine posterior estimate p(θ D) for unknown parameters θ by repeatedly applying Bayes' rule.

Example 2.3 (Reinforcement learning): RL is the ML paradigm concerned with training optimal goal-directed decision-makers (i.e., agents) through sequential interactions with an uncertain environment.

Additional examples include nonmonotonic reasoning, hard-coded algorithms (Turing machines), and probabilistic next token prediction.

The paper identifies several risks to construct validity in reasoning evaluation:

  1. A process and its product should not be conflated — reasoning is a process, not an output, and accurate QA is not sufficient for demonstrating reasoning.

  2. Chain-of-thought (CoT) traces are not trustworthy explanations — CoT is neither necessary nor sufficient for obtaining trustworthy explanations and is not guaranteed to be faithful to the model's internal decision-making.

  3. Evaluations should disentangle reasoning from recall — many benchmarks are easily gamed by instilling near-unlimited priors and experience through large-scale pre- and post-training.

Objection 1 (Def. 2.4 violates the bitter lesson; symbolic AI has failed; scaling is all you need): The paper rebuts that Points (1) and (2) are false, and (3) is speculative, noting that Def. 2.4 does not require hardcoding nor injection of human domain expertise, it is compatible with the bitter lesson.

Objection 2 (Empirical performance matters more than theoretical guarantees): The paper rebuts Sometimes yes, sometimes no, arguing that many scientifically and economically important use cases require validity, including many decision-making systems with safety or fairness implications.

The paper's contributions are: (1) An operational definition for reasoning as a learnable, rule-governed process expressed in natural language, mathematical notation, and pseudocode (Algorithm 1); and (2) Recommendations for scientific communication via a checklist of community guidelines (Appendix A).

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:


  • Improvement: Implement a dual-layer architecture where a neural network (for rule selection) is coupled with a symbolic executor (for rule application). The symbolic layer strictly enforces the formal definition of validity (Def. 2.7) by checking that every state transition S t to S t+1 is the result of an exact, properly-typed rule application from the current rule set R t. This is not a soft prompt; it's a hard computational constraint.

  • What the improved system can do: It can guarantee that any output labeled as reasoned is provably derived from a sequence of valid rule applications. It can generate a machine-checkable proof trace (like a Lean proof) for each reasoning step, making the process auditable and eliminating reasoning zombies that merely emulate reasoning through pattern matching.

  • Improvement: Build a configuration module that forces any reasoning task to be explicitly mapped to the five core components of Def. 2.4: Process, Rules, Beliefs, Evidence, State. The system will reject a task definition if any component is missing or ambiguous. It will also require a justification for any absent component (e.g., evidence = ∅ for pure deduction).

  • What the improved system can do: It can automatically assess whether a given benchmark or evaluation protocol has construct validity (Thesis 1). It can flag when a benchmark conflates process with product (e.g., QA accuracy as a proxy for reasoning) and suggest alternative metrics that measure the mechanism of reasoning, not just the final answer.

  • Improvement: Implement a meta-rule layer R M that can update the local rule set R L on-the-fly based on new evidence or performance feedback. This is not just parameter fine-tuning; it's a structured mechanism for adding, removing, or revising explicit rules (e.g., updating a theorem, changing a decision boundary, or retracting a belief). This aligns with Claim 2.5 (rules are learnable and defeasible).

  • What the improved system can do: It can adapt to distribution shifts by revising its own reasoning rules, rather than just adjusting weights. For example, if a new mathematical axiom is introduced, the system can explicitly add it to its rule set and immediately use it in subsequent reasoning steps, without retraining the entire network. It can also backtrack and retract conclusions when a rule is found to be faulty, mimicking nonmonotonic reasoning.

  • Improvement: Add a reporting layer that separately tracks and reports (a) validity (exact rule application) and (b) soundness (truth of premises relative to external ground truth). The system will never conflate the two. It will explicitly state when a conclusion is valid but unsound (e.g., The reasoning process was valid, but the initial premise was false).

  • What the improved system can do: It can provide nuanced, trustworthy outputs in high-stakes domains (e.g., medical diagnosis, legal reasoning). It can tell the user why a conclusion might be wrong (unsound premise) versus how it was derived (valid rule application). This directly addresses the trust bottleneck and prevents the illusion of correctness from a flawed but well-executed process.

  • Improvement: Implement an explicit, persistent state buffer S t = B t, E t, R t that is stored and updated at each reasoning step, rather than relying on the model's limited context window. This includes a mechanism for retrieving prior beliefs and evidence from a long-term memory store (Claim 2.8).

  • What the improved system can do: It can perform multi-step, multi-day reasoning tasks (e.g., scientific research, complex project planning) without losing track of intermediate conclusions. It can revisit and revise earlier beliefs in light of new evidence, maintaining a coherent and evolving world model. This is a step toward AGI's requirement for persistent state, as highlighted in the paper.

  • Improvement: Replace or augment chain-of-thought (CoT) explanations with a structured rule trace that shows exactly which rule was applied at each step, what the input beliefs and evidence were, and what the output belief is. This trace is guaranteed to be faithful to the internal process because it is generated by the symbolic executor, not by the language model.

  • What the improved system can do: It can provide explanations that are theoretically guaranteed to reflect the actual reasoning mechanism (Claim 2.6). This eliminates the problem of unfaithful CoT (where the model says one thing but thinks another). Users can verify each step against a known rule set, fostering epistemic trust.

  • Improvement: Add a pre-deployment module that checks whether the system's rule set and reasoning type are aligned with the deployment context (Claim D.2). For example, if the system is a probabilistic next-token predictor (Example B.3), it will refuse to be used for formal mathematical theorem proving unless a new, sound rule set is loaded.

  • What the improved system can do: It can prevent catastrophic misuse by clearly stating its operational boundaries. It can say, I am a valid probabilistic reasoner, but I am not a valid formal reasoner for this task. Please load the appropriate rule set or use a different system. This is a critical safety feature for high-stakes applications.

Summary of Capabilities: The improved AI system is not just a better language model; it is a verifiable reasoning engine. It can prove its own reasoning steps, adapt its rules, maintain long-term memory, and clearly communicate the difference between a valid process and a sound conclusion. It is designed to be trustworthy, auditable, and safe for deployment in domains where a mistake is costly.

Abstract

Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic AI, recent advances have mainly emerged from deep probabilistic generative models. Despite immense interest and rapid progress, the generative AI community has not clearly converged on operational definitions for reasoning and often implicitly rejects the historical treatment of this topic in logic and verifiable automated reasoning. This position contends that definitional ambiguity leaves the construct validity of reasoning evaluation unverifiable, undermining quantifiable progress toward trustworthy autonomous reasoning. We also contend that this ambiguity is addressable. To that end, we provide (1) operational definitions based on a synthesis of the literature, positioning valid and sound reasoning as a learnable rule-based process; and (2) a checklist for best practices in the communication of AI reasoning research.

Sources

Related papers