Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning
cs.CL, cs.AI
Submitted: 2026-06-02
Updated: 2026-08-30
Code: https://github.com/Andree-9/ACTS
Terminology
Sources
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
- Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models
- SmartThinker: Learning to Compress and Preserve Reasoning by Step-Level Length Control
- Understanding R1-Zero-Like Training: A Critical Perspective
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
- FlashThink: An Early Exit Method For Efficient Reasoning
- Reasoning Models Can Be Effective Without Thinking
- Steering LLM Thinking with Budget Guidance
- Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
- Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Chain of Draft: Thinking Faster by Writing Less
- ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
- Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
- Qwen3 Technical Report
- Understanding Reasoning in Thinking Language Models via Steering Vectors
- Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
- BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering