SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning
cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 8 pages,3 figures. Accepted at IJCNN 2026
Code: https://github.com/linqinhong/SKIP
License: http://creativecommons.org/licenses/by/4.0/
The gist: While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference latency, and even degraded performance in large
Terminology
Abstract
While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference latency, and even degraded performance in large language models (LLMs). Existing concise reasoning frameworks significantly compromise accuracy while compressing the length of output. In this paper, we propose SKIP, a self-knowledge-guided step-wise preference learning framework. Starting with lightweight fine-tuning to adjust the model's output style, SKIP introduces a carefully designed knowledge probing mechanism to guide model to output an answer at each reasoning step. Based on the correctness of intermediate steps, we construct preference data that guide the model toward more efficient and correct reasoning by leveraging DPO. Experimental results demonstrate that our method effectively improves reasoning compression while mitigating performance degradation after fine-tuning. Besides, SKIP shows strong generalization ability on out-of-distribution datasets. We further conducted ablation studies on the component parameters of our framework.
Sources
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
- OpenAI o1 System Card
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Solving math word problems with process- and outcome-based feedback
- Token-Budget-Aware LLM Reasoning
- Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
- Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
- Self-Training Elicits Concise Reasoning in Large Language Models
- Optimizing Length Compression in Large Reasoning Models
- DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models
- Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- The Llama 3 Herd of Models
- Qwen3 Technical Report
- Training Verifiers to Solve Math Word Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection