CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning
cs.AI
Submitted: 2026-07-22
Updated: 2026-09-02
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: Quantized small reasoning models can enter repetitive or otherwise unproductive trajectories, yet standard decoding does not adapt to the trajectory as it unfolds.
Terminology
Abstract
Quantized small reasoning models can enter repetitive or otherwise unproductive trajectories, yet standard decoding does not adapt to the trajectory as it unfolds. We study MGT-B, a fixed, weight-preserving controller that converts overlapping windows of uncertainty, repetition, and local-change features into position-conditional empirical tail probabilities. It accumulates mixture betting factors with a CUSUM-shaped reset, and, after an alarm, restores a coherent earlier token and key-value-cache state before constrained re-decoding. On MATH-500, a paired three-seed evaluation over 1,500 generations per method raises exact-normalized accuracy from 54.73% for vanilla decoding to 56.40% (+1.67 percentage points; problem-clustered bootstrap 95% CI [+0.47, +2.80]), while a prospectively profiled random-intervention control reaches 54.60%. The gain is positive in all three seeds and costs 5.14% more sampled tokens. Seed-0 ablations show that rollback alone does not explain the result and that an isolated repetition penalty is harmful. Five-sample self-consistency reaches 70.0% but uses about 4.84x as many tokens as MGT-B. On the harder, non-overlapping Omni-MATH evaluation, however, MGT-B obtains 16.60% versus 16.67% for vanilla (-0.07 points; clustered 95% CI [-0.33, +0.20]) with 2.10% more sampled tokens. Thus, MGT-B provides a modest, reproducible local improvement on MATH-500 in the studied configuration, but the effect does not transfer to Omni-MATH and should not be interpreted as a general improvement in mathematical reasoning.
Sources
- The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Circular Reasoning: Understanding Self-Reinforcing Loops in Large Reasoning Models
- Deep Think with Confidence
- Mitigating Overthinking in Large Reasoning Language Models via Reasoning Path Deviation Monitoring
- Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning
- Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
- Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
- Adaptive Conformal Inference by Betting
- WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
- Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection