Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning".
Jane: The gist: standard supervised fine-tuning can rapidly suppress valid reasoning traces, even when final-answer performance is partially preserved, necessitating structural reasoning reliability metrics alongside accuracy.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we’re talking about this paper called "Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning." It sounds a bit technical, but basically, it tackles a problem where models that were good at showing their work—producing those intermediate reasoning steps—start losing that ability when you adapt them to new tasks.
Jane: Right. The core issue they’re pointing out is that when you fine-tune an AI model on standard instruction-response data, which usually doesn't have those detailed reasoning traces anymore, the model can start producing plausible final answers while completely dropping the explicit steps it used to get there.
Lu: It’s about how standard fine-tuning can rapidly suppress valid reasoning traces even when you still get decent final answer performance. This paper introduces a way to measure this loss structurally, separating what is correct from what the model actually wrote during its thinking process.
Meng: So, instead of just looking at if the final answer is right or wrong, they are focusing on whether the model actually produced a complete and valid sequence of reasoning steps during adaptation. That’s a pretty important distinction for understanding why models behave this way in real-world applications.
Lalam: It’s like when you teach someone to solve a math problem step-by-step, but then you only let them see the final answer on the test and they forget how to do the steps themselves. They lose that ability to show their work, even if they still get the right final number.
The paper's summary: Tom: Okay, so what is this structural evaluation framework they’re proposing? It seems like the authors are trying to fix a problem where answer-only metrics are misleading because a model can trick you into thinking it’s still reasoning well when it isn't.
Jane: They define reasoning-trace collapse as the progressive loss of a model's ability to produce complete, non-empty, structurally valid reasoning traces during fine-tuning. They measure things like valid, empty, missing, and truncated traces separately from the final task performance.
Lu: The framework links these structural outcomes to something called reasoning-conditioned pass@one. This lets them tell the difference between a model that’s just failing because its reasoning is bad versus one that simply stops producing valid reasoning often enough during adaptation.
Meng: That separation is key because it tells us whether we need to focus on fixing the quality of the thinking process itself, or if we are dealing with a model that has simply decided not to show its work anymore.
Lalam: It’s like checking if someone is writing a coherent essay versus just having a few random words appear on the page. The paper is saying you have to check both things when you test these fine-tuned models.
The paper's improvements: Tom: So what are the actual solutions or suggestions they offer? They aren't just pointing out the problem; they’re suggesting ways to check for it and potentially fix it.
Jane: One of their major findings is that simple masking strategies can substantially preserve explicit reasoning behavior without needing those teacher-generated reasoning traces. This means you don't always need a massive amount of pre-existing reasoning data to keep the traces intact.
Lu: They also found that performance conditioned on valid reasoning can remain high even when the rate of valid reasoning falls sharply, which shows that answer accuracy alone isn't enough to guarantee good modeling behavior anymore.
Meng: That’s interesting because it means we might be able to adapt models using less expensive data or simpler methods if we focus on these structural metrics instead of just chasing the highest final accuracy score.
Lalam: So, simple masking is a practical thing you can try right away, and it seems effective at keeping the model from collapsing its reasoning ability during adaptation.
Conclusion: Tom: So to wrap up on "Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning," the main point is that we need structural reasoning reliability metrics alongside final answer accuracy, especially when you’re fine-tuning on data that doesn't have explicit traces.
Jane: They show that standard supervised fine-tuning can rapidly suppress those valid reasoning traces, and answer-only metrics really do obscure this failure mode. We should be looking at how the model produces valid traces, not just its final score.
Lu: It highlights that the representation of missing reasoning has a large effect on this collapse across different models; for instance, one model collapses Chemistry reasoning to zero percent when using a default no-think setting, while another preserves seventy-one percent with an empty-think setting.
Meng: From an engineering standpoint, we see that while simple masking strategies work well and keep performance high, they aren't a universal fix; distillation isn't uniformly better for every model and task. You have to test mitigation per model.
Lalam: I think what this means for us is that when we adapt models to new data, we can’t just assume the old reasoning skills are carried over; we need to check if those steps actually survive the fine-tuning process.
Tom: Exactly. So, it’s about moving beyond just checking if the answer is correct and demanding a structural look at how the model arrived at that answer when it’s been adapted to something new. We’ll be talking about more on how this affects agent planning in our next show.
King’s College London
cs.LG
Submitted: 2026-05-20
Updated: 2026-09-28
Comments: 22 pages, 3 tables, 3 figures. Accepted to NeurIPS 2026
Code: https://github.com/itsluketwist/thinkpack
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 88/100
The gist: The gist: standard supervised fine-tuning can rapidly suppress valid reasoning traces, even when final-answer performance is partially preserved, necessitating structural reasoning reliability
Key concepts
- Reasoning-Trace Collapse
- This is the progressive loss of a model's capacity to generate full, structurally sound reasoning steps while being fine-tuned. It occurs when models are trained on data that doesn't show explicit reasoning traces, causing them to stop producing detailed thoughts during adaptation.
- Answer Correctness vs. Reasoning-Trace Validity
- The framework separates whether a model is correct on the final answer from whether the generated reasoning trace is structurally valid (complete, empty, missing, or truncated). This distinction helps researchers see if a model fails because its thinking is bad or because it stops thinking altogether.
- Reasoning-Conditioned Pass@1
- This metric relates structural outcomes to task performance. It measures how well the model performs on a task specifically conditioned on the validity of its generated reasoning trace. This helps determine if performance drops are due to poor reasoning quality or simply a failure to produce valid traces.
- THINKPACK
- THINKPACK is an open-source library designed for evaluating models across different settings. It provides tools for constructing chats, parsing reasoning traces, computing structural metrics like valid/empty/missing rates, and masking losses during fine-tuning.
Terminology
Summary
The gist: standard supervised fine-tuning can rapidly suppress valid reasoning traces, even when final-answer performance is partially preserved, necessitating structural reasoning reliability metrics alongside accuracy.
Reasoning-Trace Collapse Definition and Framework
Reasoning-trace collapse is defined as the progressive loss of a model’s ability to produce complete, non-empty, structurally valid reasoning traces during fine-tuning
The framework introduced separates answer correctness from reasoning-trace validity
by measuring whether generated traces are valid, empty, missing, or truncated,
and relates these structural outcomes to task performance through reasoning-conditioned pass@1
. This allows researchers to distinguish between failures where reasoning is ineffective and those where the model simply stops producing valid reasoning often enough.
Impact of Fine-Tuning Data Mismatch
The mismatch between reasoning-aware model behavior and standard downstream adaptation data leads to this failure mode because many datasets used for model customisation do not contain explicit reasoning traces
. When trained on such data, a model can minimise its loss by learning to produce only the final answer, effectively treating the absence of reasoning as the desired behaviour
. This risk is particularly acute during fine-tuning because answer-only metrics can substantially obscure this failure
.
Experimental Findings on Collapse and Mitigation
The study found that ordinary fine-tuning can rapidly suppress valid reasoning traces, even when final-answer performance is partially preserved
. In several settings, performance conditioned on valid reasoning remains high while the rate of valid reasoning falls sharply,
indicating that answer accuracy alone can hide the loss of explicit reasoning behavior. When testing mitigation strategies, simple masking strategies can substantially preserve valid reasoning behaviour while maintaining strong task performance
. However, distillation is not uniformly better; for instance, it performs poorly for Olmo-3-7B on GSM8K with a final VR of 22%.
Model and Format Dependencies
The severity of collapse depends heavily on the representation of missing reasoning during fine-tuning; specifically, the representation of missing reasoning has a large effect on reasoningtrace collapse across models
. For example, for Llama-R1-8B, the default no-think setting collapses Chemistry reasoning to 0% VR, whereas empty-think preserves 71% VR. Furthermore, model-default formatting is not reliably protective,
as Qwen3-8B is more stable when reasoning tags are omitted while Olmo-3-7B and Llama-R1-8B preserve reasoning better when empty reasoning tags are included.
Structural Evaluation Metrics and Tooling
The paper introduces THINKPACK, a lightweight library for applying this evaluation framework across models,
which provides model-agnostic utilities for chat construction, trace parsing, metric computation, and loss masking
. The structural metrics reported include the valid reasoning rate (VR), empty reasoning rate (ER), missing reasoning rate (MR), and truncated reasoning rate (TR),
alongside the task-level Rpass@1. These metrics allow researchers to distinguish between performance drops due to poor reasoning quality versus a failure to produce valid traces.
Conclusion on Evaluation Practice
The findings suggest that evaluations of fine-tuned reasoning models should report structural reasoning reliability metrics in addition to final answer performance, especially when adaptation data does not contain explicit reasoning traces
. While masking is a practical mitigation, the paper concludes that basic distillation is not uniformly better,
implying that mitigation strategies must be evaluated per model rather than assumed to transfer across all scenarios.
Limitations and Robustness Checks
The study included ablations showing that collapse is not specific to the science question-answering setting, as Qwen3-8B fine-tuned on math or code datasets also exhibits reasoning trace collapse. Additionally, the analysis showed that the surviving valid traces generally remain substantively correct, rather than merely satisfying the required output format
. However, limitations exist regarding external validity and internal validity due to fixed design choices like using LoRA and greedy decoding. The paper concludes by stressing that final-answer metrics can hide reasoning-trace collapse,
emphasizing the need for structural tracking.
Appendix Details
THINKPACK is released as an open-source Python package available on PyPI and GitHub, exposing modules for chat construction, trace parsing, statistics computation, and loss masking. The implementation details confirm that the structural categories are reliably captured under the models used in this study with 100% agreement. The empirical study utilized four representative open-weight reasoning models across different training pipelines and reasoning formats to test these hypotheses. The evaluation involved fine-tuning on the Chemistry L-3 subset of SciKnowEval, GSM8K for math, and EvalPlus for code generation. The results are averaged over two random seeds, with evaluations performed every 100 optimization steps during the training run. The final metrics are reported in Tables 2, 3 and 4, distinguishing final reasoning retention from peak in-domain performance across various strategies.
Improvements for AI systems
- Bold header: Structural Reasoning Reliability Reporting
The improved system will report structural reasoning reliability metrics alongside accuracy, especially after further adaptation,
addressing the gap where answer-only evaluation can make a fine-tuned reasoning model appear functional even when explicit reasoning has largely disappeared.
- Bold header: Fine-Tuning Format Sensitivity Analysis
The system will incorporate analysis of how the representation of missing reasoning strongly affects the severity and form of collapse,
specifically comparing empty-think
versus no-think
formats to determine which is more protective against loss, as demonstrated by results showing that model-default formatting is not reliably protective.
- Bold header: Reasoning-Aware Loss Masking for Adaptation
The system will utilize simple masking strategies can substantially preserve valid reasoning behaviour while maintaining strong task performance
during fine-tuning on data without explicit traces, providing a practical alternative to teacher distillation when traces are unavailable or costly.
Abstract
Explicit reasoning models are trained to produce intermediate reasoning traces before final answers, but downstream fine-tuning is often performed on ordinary instruction--response data that contains no such traces. We show that this mismatch can induce reasoning-trace collapse: a fine-tuned model continues to produce plausible final answers while losing the structurally valid explicit reasoning traces that made it a reasoning model in the first place. We introduce a structural evaluation framework that separates answer correctness from reasoning-trace validity, measuring valid, empty, missing, and truncated reasoning alongside reasoning-conditioned task performance. Using this framework, we study four open-weight reasoning models and find that standard supervised fine-tuning can rapidly suppress valid reasoning traces, and that answer-only metrics can substantially obscure this failure: in several settings, performance conditional on valid reasoning remains high while the rate of valid reasoning falls sharply. We further show that simple loss-masking strategies can substantially mitigate collapse without requiring teacher-generated reasoning traces. These results suggest that evaluations of fine-tuned reasoning models should report structural reasoning reliability metrics in addition to final-answer performance, especially when adaptation data does not contain explicit reasoning traces.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks