Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

summary

Video file (mp4)

The gist

This paper investigates the faithfulness of chemical chain-of-thought (CoT) reasoning in large language models, revealing that these rationales often function as a "hallucination-prone molecular

In short

Research shows AI models can provide correct molecular answers while making false claims in their reasoning. These errors, called extrinsic reasoning fabrications, suggest text acts as a "molecular scratchpad" for computation. To fix this, researchers developed Chem-R-Faithful using GRPO to reward models for both accuracy and truthful reasoning.

Key concepts

Extrinsic Reasoning Fabrications (ER)
These are errors where an AI model provides a correct final answer but includes false claims within its reasoning process. For example, the Chem-R model can get a molecule exactly right while still including at least one fabricated claim about its chemical structure in its explanation.
Molecular Scratchpad
This concept suggests that an AI's verbal reasoning acts as an internal mental workspace for computation. While the text might be misleading or 'performative,' small structural snippets like SMILES fragments within the reasoning are what actually perform the heavy lifting to reach a correct conclusion.
Chem-R-Faithful
A training checkpoint developed to increase model honesty using a method called GRPO. It changes the reward system so that models only receive credit for a correct answer if their reasoning trace is also clean and free of fabrications, making the models more verifiable for scientific research.

Terminology used across episodes

This episode discusses

The paper

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad · Read on arXiv

Blue Whale Lab, National University of Singapore · Hong Kong Polytechnic University · Fudan University

Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely decoupled from answer correctness: correct answers often coexist with fabricated structural claims absent from the relevant molecules. Yet this does not make the reasoning trace computationally irrelevant. Attribution analyses suggest a shared scratchpad function expressed in model-specific forms: Chem-R and ether-0 rely on fragmented SMILES drafts, whereas ChemDFM-R emphasizes scaffold, positional, and naming cues. Notably, perturbing Chem-R's SMILES sketches degrades generation, showing that structural drafts can be causally load-bearing even when verbal structural claims are largely inert. Together, these results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad. This finding cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad".

Jane: The paper was written by Jiatong Li, Yuxuan Ren, Weida Wang, Xiao-yong Wei and Yatao Bian from Blue Whale Lab, National University of Singapore and Hong Kong Polytechnic University and Fudan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Jane, we really need to discuss this new paper titled 'Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad'. The title alone is enough to grab your attention.

Jane: It certainly is, Tom, and it sets a very serious tone for what Jiatong Li and the team from the National University of Singapore and Fudan University have uncovered. They're essentially telling us that when these models show their reasoning in chemistry, we shouldn't take those words at face value.

Tom: That sounds like a massive warning for anyone relying on AI to explain scientific processes.

Jane: It is, because the researchers suggest the text might just be a way for the model to scribble notes to itself rather than providing a real explanation.

Lu: I find that concept incredibly exciting from a research standpoint, though. If the text is acting as a mental workspace, we could potentially build entirely new architectures that optimize how AI uses these internal scratchpads to think through complex problems.

Meng: That's an interesting thought, Lu, but I'm thinking about the practical reliability for engineers. If the reasoning isn't a faithful explanation, how can we ever trust the chemical structures these models output in a real-world laboratory setting?

Jane: That's exactly what the paper is trying to address, Meng. They want to know if there's a disconnect between what the model says and what it actually does.

Lalam: This discovery could fundamentally shift our cultural expectations of AI transparency. We might have to move away from expecting logical prose and instead learn how to audit these fragmented, non-verbal workspaces to ensure they are actually grounded in reality.

Tom: It seems like we're moving from a world of "reading" AI thoughts to "deciphering" them.

Jane: Precisely, and that brings us directly to the specific ways these models are failing their own explanations.

Summary: Tom: Now that we've touched on the concept, let's look at what they actually found in 'Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad'. The team discovered that a model can give a perfectly correct molecular answer while telling you something completely false about its structure.

Jane: They call these errors "extrinsic reasoning fabrications," or ER for short, and they're quite common. For example, the Chem-R model gets the molecule exactly right about thirteen percent of the time even though it includes at least one fabricated claim in its reasoning trace.

Meng: That sounds like a nightmare for anyone trying to debug these systems. How do you even begin to fix a model that is simultaneously successful in its task but totally dishonest about its process?

Lu: You have to look past the words, Meng. The researchers found that if you change the verbal claims in the reasoning, it barely affects the answer, but if you mess with those tiny SMILES fragments they scribble down, the whole thing falls apart.

Tom: So those little structural snippets are actually doing the heavy lifting?

Lu: Exactly, which means the text is mostly just a distraction while those fragments act as a causal scratchpad for the actual computation.

Jane: It really shows that we've been overestimating how much these models actually "understand" the language they're using to explain themselves.

Lalam: This creates a risk of what I'd call "performative reasoning." We might see AI systems that look incredibly intelligent and logical to a human observer, but are actually just performing a series of disconnected steps that don't actually support their conclusions.

Meng: That kind of mismatch could lead to serious safety issues in high-stakes fields like drug discovery if we aren't careful.

Tom: It definitely makes you wonder if there is any way to actually force these models to be honest.

Improvements: Tom: We've seen how messy the reasoning can get, so let's talk about how the researchers suggest we fix it in 'Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad'. They didn't just identify the problem; they actually developed a way to train models to be more faithful.

Jane: They introduced a checkpoint called Chem-R-Faithful using a method called GRPO. Essentially, they changed the reward system so that the model only gets credit for a correct answer if its reasoning trace is also clean and free of fabrications.

Meng: I'm curious about the trade-offs there. Did they have to sacrifice a lot of accuracy just to make the models more honest?

Jane: Actually, Meng, they managed to significantly reduce those fabrications while actually preserving or even improving their overall performance.

Lu: This is such a brilliant move toward process-level supervision. Instead of just checking if the student got the right answer on a test, we're finally checking if they actually followed the correct steps to get there!

Tom: That would change how we approach training scientific AI entirely, wouldn't it?

Lu: It really would, because it allows us to build models that are verifiable rather than just predictive.

Meng: Implementing that kind of real-time verification in a production environment sounds like a massive engineering challenge.

Jane: It certainly is more complex to set up, but the results show it's a necessary step for reliable science.

Lalam: This shift could move our relationship with AI from one of blind trust to one of active, collaborative verification. We won't just be accepting answers; we'll be participating in a reasoning process that we can actually audit and understand.

Tom: It sounds like they've given us a real way forward for making these tools useful for actual research.

Conclusion: Tom: We have covered so much ground today, from the discovery of these "molecular scratchpads" to the potential of process-level supervision.

Jane: It's been a fascinating look at how 'Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad' challenges our assumptions about AI reasoning.

Lu: I'm walking away thinking about how we can design even better mental workspaces for these models to inhabit.

Meng: And I'll be thinking about how we can build the robust verification pipelines needed to make this a reality in the lab.

Lalam: Ultimately, this research pushes us toward a future where our digital tools are built on a foundation of verifiable truth.

Tom: Thanks to everyone for joining the discussion, and thanks to our team for their insights. We'll see you next time!

More episodes

← Home