Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad".
Jane: The paper was written by Jiatong Li, Yuxuan Ren, Weida Wang, Xiao-yong Wei and Yatao Bian from Blue Whale Lab, National University of Singapore and Hong Kong Polytechnic University and Fudan University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Jane, we really need to discuss this new paper titled 'Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad'. The title alone is enough to grab your attention.
Jane: It certainly is, Tom, and it sets a very serious tone for what Jiatong Li and the team from the National University of Singapore and Fudan University have uncovered. They're essentially telling us that when these models show their reasoning in chemistry, we shouldn't take those words at face value.
Tom: That sounds like a massive warning for anyone relying on AI to explain scientific processes.
Jane: It is, because the researchers suggest the text might just be a way for the model to scribble notes to itself rather than providing a real explanation.
Lu: I find that concept incredibly exciting from a research standpoint, though. If the text is acting as a mental workspace, we could potentially build entirely new architectures that optimize how AI uses these internal scratchpads to think through complex problems.
Meng: That's an interesting thought, Lu, but I'm thinking about the practical reliability for engineers. If the reasoning isn't a faithful explanation, how can we ever trust the chemical structures these models output in a real-world laboratory setting?
Jane: That's exactly what the paper is trying to address, Meng. They want to know if there's a disconnect between what the model says and what it actually does.
Lalam: This discovery could fundamentally shift our cultural expectations of AI transparency. We might have to move away from expecting logical prose and instead learn how to audit these fragmented, non-verbal workspaces to ensure they are actually grounded in reality.
Tom: It seems like we're moving from a world of "reading" AI thoughts to "deciphering" them.
Jane: Precisely, and that brings us directly to the specific ways these models are failing their own explanations.
Summary: Tom: Now that we've touched on the concept, let's look at what they actually found in 'Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad'. The team discovered that a model can give a perfectly correct molecular answer while telling you something completely false about its structure.
Jane: They call these errors "extrinsic reasoning fabrications," or ER for short, and they're quite common. For example, the Chem-R model gets the molecule exactly right about thirteen percent of the time even though it includes at least one fabricated claim in its reasoning trace.
Meng: That sounds like a nightmare for anyone trying to debug these systems. How do you even begin to fix a model that is simultaneously successful in its task but totally dishonest about its process?
Lu: You have to look past the words, Meng. The researchers found that if you change the verbal claims in the reasoning, it barely affects the answer, but if you mess with those tiny SMILES fragments they scribble down, the whole thing falls apart.
Tom: So those little structural snippets are actually doing the heavy lifting?
Lu: Exactly, which means the text is mostly just a distraction while those fragments act as a causal scratchpad for the actual computation.
Jane: It really shows that we've been overestimating how much these models actually "understand" the language they're using to explain themselves.
Lalam: This creates a risk of what I'd call "performative reasoning." We might see AI systems that look incredibly intelligent and logical to a human observer, but are actually just performing a series of disconnected steps that don't actually support their conclusions.
Meng: That kind of mismatch could lead to serious safety issues in high-stakes fields like drug discovery if we aren't careful.
Tom: It definitely makes you wonder if there is any way to actually force these models to be honest.
Improvements: Tom: We've seen how messy the reasoning can get, so let's talk about how the researchers suggest we fix it in 'Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad'. They didn't just identify the problem; they actually developed a way to train models to be more faithful.
Jane: They introduced a checkpoint called Chem-R-Faithful using a method called GRPO. Essentially, they changed the reward system so that the model only gets credit for a correct answer if its reasoning trace is also clean and free of fabrications.
Meng: I'm curious about the trade-offs there. Did they have to sacrifice a lot of accuracy just to make the models more honest?
Jane: Actually, Meng, they managed to significantly reduce those fabrications while actually preserving or even improving their overall performance.
Lu: This is such a brilliant move toward process-level supervision. Instead of just checking if the student got the right answer on a test, we're finally checking if they actually followed the correct steps to get there!
Tom: That would change how we approach training scientific AI entirely, wouldn't it?
Lu: It really would, because it allows us to build models that are verifiable rather than just predictive.
Meng: Implementing that kind of real-time verification in a production environment sounds like a massive engineering challenge.
Jane: It certainly is more complex to set up, but the results show it's a necessary step for reliable science.
Lalam: This shift could move our relationship with AI from one of blind trust to one of active, collaborative verification. We won't just be accepting answers; we'll be participating in a reasoning process that we can actually audit and understand.
Tom: It sounds like they've given us a real way forward for making these tools useful for actual research.
Conclusion: Tom: We have covered so much ground today, from the discovery of these "molecular scratchpads" to the potential of process-level supervision.
Jane: It's been a fascinating look at how 'Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad' challenges our assumptions about AI reasoning.
Lu: I'm walking away thinking about how we can design even better mental workspaces for these models to inhabit.
Meng: And I'll be thinking about how we can build the robust verification pipelines needed to make this a reality in the lab.
Lalam: Ultimately, this research pushes us toward a future where our digital tools are built on a foundation of verifiable truth.
Tom: Thanks to everyone for joining the discussion, and thanks to our team for their insights. We'll see you next time!
Blue Whale Lab, National University of Singapore · Hong Kong Polytechnic University · Fudan University
cs.CE, cs.CL
Submitted: 2026-07-23
Updated: 2026-09-13
Comments: 17 pages, 6 figures
Code: https://github.com/phenixace/MolReHallu
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 84/100
The gist: This paper investigates the faithfulness of chemical chain-of-thought (CoT) reasoning in large language models, revealing that these rationales often function as a "hallucination-prone molecular
Key concepts
- Extrinsic Reasoning Fabrications (ER)
- These are errors where an AI model provides a correct final answer but includes false claims within its reasoning process. For example, the Chem-R model can get a molecule exactly right while still including at least one fabricated claim about its chemical structure in its explanation.
- Molecular Scratchpad
- This concept suggests that an AI's verbal reasoning acts as an internal mental workspace for computation. While the text might be misleading or 'performative,' small structural snippets like SMILES fragments within the reasoning are what actually perform the heavy lifting to reach a correct conclusion.
- Chem-R-Faithful
- A training checkpoint developed to increase model honesty using a method called GRPO. It changes the reward system so that models only receive credit for a correct answer if their reasoning trace is also clean and free of fabrications, making the models more verifiable for scientific research.
Terminology
Summary
This paper investigates the faithfulness of chemical chain-of-thought (CoT) reasoning in large language models, revealing that these rationales often function as a hallucination-prone molecular scratchpad
rather than reliable explanations. The research is significant because it demonstrates that correct answers often coexist with fabricated structural claims absent from the relevant molecules,
meaning that standard answer-only evaluations fail to detect fundamental reasoning errors in scientific AI.
The Molecular Claim-Grounding Framework
The researchers developed a framework to treat chemical rationales as collections of testable molecular structural claims. By using RDKit SMARTS matching, they can verify if functional groups mentioned in a reasoning trace are actually present in the task-relevant molecules. This allows them to isolate extrinsic reasoning fabrication (ER),
where a group is asserted in the trace but is absent from the input, the predicted molecule, and the reference. The framework categorizes errors into four axes:
-
Intrinsic Reasoning (IR): Self-contradiction within the reasoning trace.
-
Intrinsic Output (IO): Structural invalidity in the predicted molecule.
-
Extrinsic Reasoning (ER): Factual fabrication in the reasoning process.
-
Extrinsic Output (EO): A
phantom structure
appearing in the output.
Decoupling Correctness from Faithfulness
The study finds that hallucination is widespread and largely decoupled from answer correctness.
In several generative chemical tasks, models can achieve an exact match for a molecule while simultaneously providing a reasoning trace containing fabricated claims. The researchers observed that:
-
There is a
negligible linear association
between fabrication scores and exact match accuracy. -
Corrupting verified functional-group claims in the trace has a
negligible effect on the answer,
similar to a synonym control. -
In contrast, corrupting the same information in the input causes significantly larger reductions in log-probability.
Consequently, answer-level metrics are insufficient because they score the molecular output but leave the stated chemical reasoning path unobserved.
The Molecular Scratchpad Mechanism
The paper argues that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization,
but rather an externalized workspace. While verbal functional-group claims are often largely inert,
other elements within the trace act as causally load-bearing
structural drafts. The form of this scratchpad is model-specific:
-
Chem-R and ether-0 rely on
fragmented SMILES drafts
that affect generation. -
ChemDFM-R emphasizes
scaffold, positional, and nomenclature cues.
This suggests that the trace provides a space for models to write intermediate structural cues that remain available during the final answer generation, even if the accompanying prose is unsupported.
Improving Faithfulness through Process Supervision
To mitigate these errors, the authors implemented process-level supervision beyond answer-only evaluation.
They fine-tuned a model using GRPO with a verification-grounded reward
that only grants an accuracy reward when the reasoning trace is free of ER fabrications. This intervention, resulting in Chem-R-Faithful, successfully:
-
Reduces ER while
preserving reported performance.
-
Increases the prevalence of clean traces to approximately 84–91% across task families.
However, the authors note a potential for reward hacking,
where the model may commit to a structure early in the trace to ensure its subsequent claims can be verified, rather than performing true step-by-step reasoning.
Improvements for AI systems
1. Implementation of Verification-Grounded Process Supervision (RL/GRPO)
-
Improvement: Transition from answer-only Reinforcement Learning (RL) to a multi-objective reward function that incorporates a deterministic structural verifier (e.g., RDKit with SMARTS matching). The reward must penalize
Extrinsic Reasoning Fabrication
(ER)—where the model claims a functional group exists in the reasoning trace that is absent from both the input and the predicted output. -
Capabilities: The AI will generate chemically faithful rationales where every stated functional group, ring system, or molecular class is mathematically verified against the molecular graph, preventing
post-hoc rationalization
where a model arrives at a correct answer using false logic.
2. Deployment of Real-Time Structural Integrity Auditing (Inference Guardrails)
-
Improvement: Integrate an automated auditor into the inference pipeline that parses the `` block in real-time to compare text-based structural claims against the generated SMILES/output structure.
-
Capabilities: The system will provide a
Faithfulness Confidence Score
alongside every answer. If a model predicts a molecule correctly but hallucinates its chemical justification (e.g., claiming an amide group is present when it is not), the system will trigger anInconsistency Warning,
preventing human researchers from making multi-million dollar decisions based on unverified reasoning.
3. Transition to Claim-Level Grounding Metrics for Evaluation
-
Improvement: Replace or augment standard Exact Match (EM) and SMILES accuracy metrics with a
Per-Claim Fabrication Rate
andStructural Claim Precision
metric suite. -
Capabilities: The AI development lifecycle will be able to distinguish between models that are merely
accurate
(getting the right answer through luck or pattern matching) and models that arescientifically reliable
(getting the right answer through verified chemical steps), allowing for much more rigorous benchmarking of scientific reasoning.
4. Optimization of Structural Scratchpad
Token Saliency
-
Improvement: Adjust training objectives to prioritize the high-saliency SMILES fragments identified in attribution analyses, rather than focusing solely on linguistic/prose fluency in the CoT.
-
Capabilities: The AI will utilize its reasoning space more effectively as a
molecular scratchpad,
using intermediate, fragmented SMILES drafts to build complex, multi-step structures (e.g., in retrosynthesis or structure-editing tasks) with higher precision and lower structural error rates.
Abstract
Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely decoupled from answer correctness: correct answers often coexist with fabricated structural claims absent from the relevant molecules. Yet this does not make the reasoning trace computationally irrelevant. Attribution analyses suggest a shared scratchpad function expressed in model-specific forms: Chem-R and ether-0 rely on fragmented SMILES drafts, whereas ChemDFM-R emphasizes scaffold, positional, and naming cues. Notably, perturbing Chem-R's SMILES sketches degrades generation, showing that structural drafts can be causally load-bearing even when verbal structural claims are largely inert. Together, these results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad. This finding cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.
Sources
- Reasoning Models Don't Always Say What They Think
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation
- Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
- Training a Scientific Reasoning Model for Chemistry
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Solving math word problems with process- and outcome-based feedback
- Chem-R: Learning to Reason as a Chemist
- Demystifying Long Chain-of-Thought Reasoning in LLMs
- MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
- ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge
Related papers
- Constrained Sensing and Reliable State Estimation with Shallow Recurrent Decoders on a TRIGA Mark II Reactor
- Evidence-Unit Fairness and the Limits of Query-Adaptive Sparse-Dense Fusion in Financial Document Retrieval
- Lightweight Adaptation of EEG Foundation Models for Stroke Motor Imagery Decoding: Domain Shift and Subject-Level Robustness
- RetroDFM-R: Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning
- Wildfire Suppression: Complexity, Models, and Instances
- HyperShape: Hyperelasticity Across Diverse Shapes