Can your AI agent be cheaper? Investigating the effects of task specifications on token spend in agentic coding tasks
cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/github/spec-kit
Terminology
Sources
- OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
- When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation
- Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
- How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Value of Information: A Framework for Human-Agent Communication
- SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
- AI Agents That Matter
- What Makes a Good Bug Report for an AI Agent?
- ClarEval: A Benchmark for Evaluating Clarification Skills of Code Agents under Ambiguous Instructions
- SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
- Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
- Agentless: Demystifying LLM-based Software Engineering Agents
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection