Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories
cs.LG, cs.CL
Submitted: 2026-07-07
Updated: 2026-09-01
Terminology
Sources
- Are Latent Reasoning Models Easily Interpretable?
- Training Large Language Models to Reason in a Continuous Latent Space
- How to use and interpret activation patching
- Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
- Training Verifiers to Solve Math Word Problems
- Measuring Faithfulness in Chain-of-Thought Reasoning
- How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
- Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
- LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
- Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought
- Multiple-Choice Questions are Efficient and Robust LLM Evaluators
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks