TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models
cs.CR
Submitted: 2026-03-02
Updated: 2026-09-01
Comments: 23 pages,18 figures,8 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Reasoning Models (LRMs) introduce a reasoning-level attack surface: adversaries can corrupt intermediate inferences while preserving a plausible trace and an apparently benign output.
Terminology
Abstract
Large Reasoning Models (LRMs) introduce a reasoning-level attack surface: adversaries can corrupt intermediate inferences while preserving a plausible trace and an apparently benign output. Existing output guardrails cannot reliably identify where such a trace first becomes unsupported. We present TraceGuard, a compact, locally deployable reasoning firewall that treats model-generated reasoning as untrusted input. Its design combines grounded generation of verifiable audit traces, Step-Aware Supervised Fine-Tuning (SSFT) for process-level supervision, and Verifier-Guided Reinforcement Learning (VGRL) for hardening against difficult reasoning traces. TraceGuard audits intermediate steps, localizes the initial Point of Fracture, and grounds its final decision in the complete audit evidence. We evaluate TraceGuard across heterogeneous open-weight architectures, reasoning domains, and reasoning-integrity attack families. A compact Qwen3-4B-Guard substantially outperforms an unaligned 20B model under strict end-to-end detection. Its auditing behavior transfers to attack families excluded from training, resists in-scope black-box probing, and remains robust in an additional white-box stress test. Overall, 210,456 step-level audit decisions support compact, process-aligned verification as an effective, deployable defense boundary for reasoning systems.
Sources
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs
- Constitutional AI: Harmlessness from AI Feedback
- Training Verifiers to Solve Math Word Problems
- LoRA: Low-Rank Adaptation of Large Language Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
- CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
- Solving math word problems with process- and outcome-based feedback
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs