SFAD: Speculative Factuality-Aware Decoding
cs.CL
Submitted: 2026-09-01
Updated: 2026-09-03
License: http://creativecommons.org/licenses/by/4.0/
The gist: As one of the most critical challenges in large language models, contextual faithfulness directly determines their reliability in knowledge-intensive applications.
Terminology
Abstract
As one of the most critical challenges in large language models, contextual faithfulness directly determines their reliability in knowledge-intensive applications. This task is particularly challenging as it requires balancing factual consistency with generation efficiency. Contrastive decoding methods require dual forward passes (with and without context) to compare model outputs, doubling inference computational overhead, while post-training alignment demands extensive reinforcement learning with substantial computational overhead. To address this challenge, we present SFAD, a speculative decoding framework that enhances contextual faithfulness without inference degradation. We first construct ConFide, a preference dataset with fine-grained atomic perturbations, to train a context-faithful draft model via Direct Preference Optimization. During inference, Epistemic Friction detects potential hallucinations by quantifying distributional tension weighted by specialist certainty. When friction exceeds the threshold, Asymmetric Logit Steering refines the target distribution through residual-based logit injection; otherwise, standard speculation proceeds. Extensive experiments demonstrate that SFAD substantially improves faithfulness while achieving 2.48 times speedup, offering a practical solution for efficient LLMs.
Sources
- GPT-4 Technical Report
- Accelerating Large Language Model Decoding with Speculative Sampling
- Leveraging Logical Rules in Knowledge Editing: A Cherry on the Top
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GPT-4o System Card
- Challenges and Applications of Large Language Models
- MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
- HAGRID: A Human-LLM Collaborative Dataset for Generative Information-Seeking with Attribution
- Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models
- Measuring short-form factuality in large language models
- Context-aware Decoding Reduces Hallucination in Query-focused Summarization
- Qwen3 Technical Report
- Investigating CoT Monitorability in Large Reasoning Models
- Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
- When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering