From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
cs.CL, cs.AI
Submitted: 2026-02-27
Updated: 2026-08-28
Code: https://github.com/UKPLab/arxiv2026-controllable-reasoning-models
Terminology
Sources
- Training Verifiers to Solve Math Word Problems
- Chain-of-Sanitized-Thoughts: Plugging PII Leakage in CoT of Large Reasoning Models
- Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
- Phi-4 Technical Report
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control
- ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
- Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
- gpt-oss-120b & gpt-oss-20b Model Card
- Effectively Controlling Reasoning Models through Thinking Intervention
- Qwen3 Technical Report
- Dynamic Early Exit in Reasoning Models
- Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
- Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
- Evaluating Language Model Reasoning about Confidential Information
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering