On the Token Value Inequality in Efficient Reasoning
cs.CL, cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Phi-4-reasoning Technical Report
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- Training Language Models to Reason Efficiently
- SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
- Dense Reward for Free in Reinforcement Learning from Human Feedback
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
- Learning to Route LLMs with Confidence Tokens
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- Think before you speak: Training Language Models With Pause Tokens
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Training Large Language Models to Reason in a Continuous Latent Space
- Measuring Massive Multitask Language Understanding
- Measuring Mathematical Problem Solving With the MATH Dataset
- The Curious Case of Neural Text Degeneration
- Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering