Adaptive Triggering for Bias Correction in LLM Reasoning
cs.CL, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 10 pages, 6 figures, Under review
Code: https://github.com/clairekim59/adaptive-triggering
License: http://creativecommons.org/licenses/by/4.0/
The gist: Chain-of-thought prompting can expose and amplify demographic stereotypes within an LLM's intermediate reasoning and create a failure mode that final-answer debiasing alone cannot address.
Terminology
Abstract
Chain-of-thought prompting can expose and amplify demographic stereotypes within an LLM's intermediate reasoning and create a failure mode that final-answer debiasing alone cannot address. Mitigating such bias during generation presents a fundamental timing problem: intervening too late allows biased reasoning to propagate, while unnecessarily intervening can disrupt otherwise correct reasoning. Existing approaches largely avoid this decision by either evaluating completed reasoning chains post hoc or intervening at predetermined steps, leaving open when a developing reasoning trajectory provides sufficient evidence to warrant correction. We formulate this decision as an online change-point detection problem. A per-step bias signal updates a CUSUM statistic and a targeted correction is injected only when accumulated evidence crosses a detector-specific threshold calibrated on held-out data. We instantiate the framework with a white-box signal derived from next-token probabilities and a black-box signal obtained from an LLM judge, enabling deployment with both open-weight and hosted models. On gpt-4o-mini adaptive black-box triggering recovers most of the disambiguated-context accuracy lost under fixed-interval intervention while requiring substantially fewer interventions. That result holds even with an independent judge. Across six open-weight models, the white-box signal improves ambiguous-item accuracy on all six but reduces disambiguated-item accuracy on five because it cannot distinguish unsupported stereotype reliance from correct, stereotype-congruent evidence.
Sources
- Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
- Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector
- Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
- Self-Debias: Self-correcting for Debiasing Large Language Models
- Reliable Control-Point Selection for Steering Reasoning in Large Language Models
- Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?
- Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
- CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning
- Mitigating Overthinking in Large Reasoning Language Models via Reasoning Path Deviation Monitoring
- Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
- Sequential statistical inference for Large Language Models: Representation, validity, and monitoring
- DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
- Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering