Compliance vs. Sensibility: On the Reasoning Controllability in Large Language Models
cs.CL, cs.AI
Submitted: 2026-04-29
Updated: 2026-09-06
Comments: Accepted to the Findings of EMNLP 2026
Code: https://github.com/Xingwei-Tan/compliance_sensibility
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Large Language Models (LLMs) acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicited via Chain-of-Thought (CoT).
Terminology
Abstract
Large Language Models (LLMs) acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicited via Chain-of-Thought (CoT). However, whether fundamental reasoning patterns, such as induction, deduction, and abduction, can be decoupled from specific problem instances remains a critical challenge for model controllability. In this paper, we present the first systematic investigation of this problem through the lens of reasoning conflicts, an explicit tension between parametric and contextual information induced by mandating logical schemata that deviate from those expected for a target task. Our evaluation reveals that LLMs consistently prioritize sensibility over compliance, favoring task-appropriate reasoning patterns despite conflicting instructions. We further demonstrate that reasoning conflicts are internally detectable, as confidence scores drop during conflicting episodes. Probing reveals instruction decodability even without compliance, while CKA identifies geometric differences across models and spans. Leveraging these insights, we steer models towards compliance, increasing instruction following by up to 29%. Overall, our findings establish that while LLM reasoning is anchored to concrete instances, active mechanistic interventions can effectively decouple logical schemata from data, offering a path toward improved controllability, faithfulness, and generalizability.
Sources
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Fundamental Reasoning Paradigms Induce Out-of-Domain Generalization in Language Models
- PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
- Language Models (Mostly) Know What They Know
- Are DeepSeek R1 And Other Reasoning Models More Faithful?
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning
- Evaluating the Logical Reasoning Abilities of Large Reasoning Models
- Logical Reasoning in Large Language Models: A Survey
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
- Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
- Olmo 3
- EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
- Qwen3 Technical Report
- A Comprehensive Study of Knowledge Editing for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering