The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts
cs.AI, cs.PF
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/Sachin-Wani/reasoning_tax
Terminology
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen3 Technical Report
- OckBench: Measuring the Efficiency of LLM Reasoning
- ReEfBench: Quantifying the Reasoning Efficiency of LLMs
- Measuring Massive Multitask Language Understanding
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Generalizing Verifiable Instruction Following
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- MathArena: Evaluating LLMs on Uncontaminated Math Competitions
- Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- APEX-Agents
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection