How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks".
Tom: Agentic coding tasks are uniquely expensive, consuming orders of magnitude more tokens than code reasoning and code chat tasks, and models vary substantially in token efficiency across different frontier LLMs.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about that title again, "How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks," because it perfectly captures the tension between the hype around agent capabilities and the very real economic reality of using them.
Jane: Exactly; it frames the whole discussion around a practical problem for developers and users: if we don't know how much this coding agent will cost us, we can't truly manage our expectations or budget effectively.
Lu: The authors are pulling from some major institutions, like the University of Michigan and Stanford University, which tells you this isn't just a niche study; it’s coming from a group that understands the broader implications for how AI is deployed in complex software engineering environments.
Meng: I saw that they are analyzing trajectories from eight frontier LLMs on SWE-bench Verified, and I wonder if testing them across such a diverse set of models helps isolate behavior caused by the model architecture versus the inherent difficulty of the task itself.
Lalam: The authors are opening up all agent trajectories, which is huge because it lets everyone see exactly what these agents are doing step-by-step when they spend their resources, which is a great way to build trust.
The paper's summary: Tom: So, the main takeaway from this study on "How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks" is that agentic tasks are uniquely expensive compared to simpler things like code reasoning or code chat, consuming one thousand times more tokens <ref:2604.22750#pg0,Analyzing and Predicting Token Consumption in Agentic Coding Tasks>.
Jane: That massive difference is striking because it shows that the way agents operate—reading repositories, calling tools, and iterating—demands a huge input token ratio compared to just asking a model a question directly.
Lu: The paper also found that token usage isn't stable; runs on the same task can vary by up to thirty times in total tokens, which suggests that the process itself is inherently stochastic rather than deterministic.
Meng: That variability is something I’ve seen in testing, and it makes it really difficult for us to guarantee a fixed cost for a specific type of agentic workflow beforehand.
Lalam: And they also showed that higher token usage doesn't automatically mean better accuracy; sometimes the performance peaks at an intermediate cost point before leveling off at higher costs.
The paper's improvements: Tom: The paper proposes a few key directions for improvement, and one of the big ones is formulating a pre-execution agent token consumption prediction task to benchmark models against.
Jane: They are trying to solve that prediction problem, showing that while agents can capture coarse trends in usage, they still have a significant gap when trying to predict exact instance-level usage across different models.
Lu: They specifically highlight that output token usage is easier to predict than input token usage because the context construction adds a lot of uncertainty right at the beginning of the process.
Meng: This finding about output versus input prediction is practically useful because it tells us where we can get a more reliable signal for budget alerts, even if it’s not perfect.
Lalam: I think that ability to give us that coarse-grained signal for relative cost is a big step toward making these agents feel more manageable and predictable for the end user.
Conclusion: Tom: So, wrapping up this discussion on "How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks," we see that agentic tasks are way more token-heavy than basic reasoning, and predicting exact costs before execution is still a genuinely difficult task for the current models.
Jane: It’s clear that while self-prediction offers some useful coarse trends, the inherent stochastic nature of these workflows means we can't rely on simple upfront pricing right now.
Lu: The paper successfully illuminated where tokens go in agentic coding and what we can anticipate before execution, which is a solid foundation for future research into agent economics.
Meng: For practical implementation, the biggest lesson seems to be that consumption-based pricing will remain the most straightforward option until these prediction methods become much more robust.
Lalam: I think the ability for agents to predict their own cost at least at a coarse level is a concrete signal for providers to issue early warnings, which is something we can start building on immediately.
Tom: That’s right; it confirms that understanding token consumption patterns is essential for anyone building or using these sophisticated AI systems. We'll keep an eye on how the community develops methods to make those predictions more reliable before we move on to the next piece of research.
Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea
University of Michigan · Stanford University · All Hands AI 4Google Deepmind Microsoft AI Massachusetts Institute of Technology
cs.CL, cs.AI, cs.CY, cs.HC, cs.SE
Submitted: 2026-04-24
Updated: 2026-10-02
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: Agentic coding tasks are uniquely expensive, consuming orders of magnitude more tokens than code reasoning and code chat tasks, and models vary substantially in token efficiency across different
Key concepts
- Agentic Tasks
- These are complex coding workflows where an AI agent performs multi-step actions like planning, executing code, and iterating. They are uniquely expensive because agents repeatedly feed context into models from different sources, leading to a massive input/output ratio compared to simpler tasks.
- Non-Monotonic Accuracy
- This means that as you spend more tokens (increase cost), performance does not always improve. In coding, this often happens because extra computation leads to redundant edits or exploration rather than deeper reasoning, causing accuracy to peak at a moderate cost level and then plateau.
- Self-Prediction Bias
- Agents can estimate their own token usage before running a task. However, this self-prediction is biased; agents systematically underestimate both input and output tokens. While not perfectly accurate for every instance, it provides a useful coarse signal to identify which tasks are likely to be very expensive.
- Stochastic Consumption
- Token usage for the same coding problem can vary wildly across different runs, sometimes differing by 30 times. This high variance makes predicting the exact cost before execution fundamentally difficult because agent behavior becomes unstable on more complex problems.
Terminology
Summary
Agentic coding tasks are uniquely expensive, consuming orders of magnitude more tokens than code reasoning and code chat tasks, and models vary substantially in token efficiency across different frontier LLMs. This systematic study analyzes token consumption patterns in agentic coding tasks using trajectories from eight frontier LLMs on SWE-bench Verified to address where agents spend money, which models are most efficient, and whether agents can predict their own costs before execution.
Key Findings on Token Consumption Patterns
The analysis reveals several critical insights into the economics of AI agents. First, agentic tasks are uniquely expensive,
consuming 1000× more tokens than code reasoning and code chat, with input tokens rather than output tokens driving the overall cost.
Second, token usage is highly variable and inherently stochastic: runs on the same task can differ by up to 30× in total tokens,
and higher token usage does not translate into higher accuracy; instead, accuracy often peaks at intermediate cost and saturates at higher costs.
Third, models show significant variation in efficiency: on the same tasks, Kimi-K2 and Claude Sonnet-4.5, on average, consume over 1.5 million more tokens than GPT-5.
Finally, a fundamental gap exists between human perception and computational effort: task difficulty rated by human experts only weakly aligns with actual token costs.
Token Usage Dynamics Across Tasks and Runs
The study systematically compared agentic coding against code reasoning and code chat tasks, showing that agentic workflows accumulate the information from different sources and the same context gets fed into the models repeatedly, resulting in a dramatically higher input/output ratio than the other two task types.
Furthermore, token usage exhibits high variance both across problems and repeated runs of the same problem. The most expensive problem costs ∼7 million more tokens than the cheapest,
and high-token-cost problems exhibit larger cross-run variance,
suggesting that agent behavior becomes increasingly unstable on more complex tasks.
This variability makes upfront cost prediction fundamentally difficult.
Accuracy vs. Cost Trade-offs
The research demonstrates a non-monotonic relationship between token expenditure and performance. When working on the same problem, agent performance peaks at the intermediate-cost run and then saturates with higher costs.
This is consistent with findings that additional reasoning steps or longer chains of thought do not necessarily improve accuracy,
suggesting that increased computation may reflect unproductive exploration rather than deeper reasoning.
Moreover, analyzing fine-grained actions shows that high-cost runs are associated with repeated view and edits for the same file,
indicating that these expensive runs are often dominated by redundancy
rather than productive progress.
Predicting Agent Token Consumption Before Execution
The paper formalized the task of predicting token consumption before execution, focusing on self-prediction where the agent estimates its own usage. The results indicate that agents can capture coarse trends in token consumption, but achieve only weak-to-moderate correlation with real usage across models.
Specifically, output-token usage is easier to predict than input-token usage,
reflecting the uncertainty introduced by context construction. Although accurate instance-level prediction remains challenging, self-prediction provides a useful coarse-grained signal of relative cost
that can support early budget alerts.
Model Token Efficiency and Prediction Performance
Models vary substantially in their token efficiency across tasks, with efficiency differences stemming from model-specific behavior rather than intrinsic task difficulty.
For instance, GPT-5 and GPT-5.2 achieve strong accuracy at low cost, while Kimi-K2 remains an outlier with both the highest cost and the lowest accuracy. In terms of prediction capability, Agents systematically underestimate both input and output token usage,
a bias that persists even when no in-context example is provided. While self-prediction is challenging, it offers a measurable signal for cost transparency without requiring additional infrastructure.
Implications for Pricing and Future Research
The findings highlight the tension between variable consumption and the need for predictable pricing models. Because agent trajectories are inherently stochastic,
purely upfront pricing remains difficult, suggesting that consumption-based pricing will likely stay the most practical option until pre-execution estimation becomes more reliable.
The study suggests that agents themselves can potentially serve as useful predictors of their own cost, at least at the coarse-grained level of identifying high-cost tasks,
which is sufficient for providers to issue early warnings. This work provides a foundation for future research aimed at building more controllable and transparent agent pricing schemes.
Conclusion
Self-prediction captures coarse-grained trends in token usage but remains noisy at the instance level.
Overall, the study confirms that predicting token usage before execution is a genuinely difficult task for current models,
yet it offers concrete steps toward more predictable and user-aligned agent pricing by identifying high-cost tasks.
**(Note: The provided text does not contain a section explicitly titled The gist
or a single sentence summary under that heading.
Improvements for AI systems
Here are specific improvements to AI systems based on the findings of this research, along with what those improved systems could achieve:
-
A mechanism for
Pre-Execution Cost Bidding
for Agentic Tasks: The system should not execute a long-horizon agent task (like complex coding) without first running a self-prediction sub-routine that estimates the total expected token cost and confidence level. -
Dynamic Model Selection Based on Predicted Budget: Agents should be able to query multiple frontier models and select the one that offers the best trade-off between accuracy and predicted cost, potentially switching to a more token-efficient model if the task complexity is high or if the predicted budget is exceeded.
-
Adaptive Stopping Criteria Guided by Cost Thresholds: The agent's internal planning mechanism should be augmented to monitor accumulated tokens against a user-defined or pre-set budget threshold (derived from the initial prediction). If the cost approaches this threshold, the agent should automatically trigger a
Plan to Stop
phase, prioritizing summary generation over further exploratory actions, mitigating runaway token consumption. -
Cost-Aware Task Decomposition: Before starting complex tasks, the agent should decompose them into sub-tasks and assign an estimated token budget to each phase (e.g., Exploration vs. Debugging). This allows for more granular cost tracking and prevents a single high-cost exploratory step from bankrupting the entire task's budget.
-
Transparency Layer for User Trust: Implement a user-facing interface that displays the predicted token cost and confidence score alongside the task description before execution begins, allowing users to make informed decisions about committing to an expensive run.
This improved AI system can achieve:
-
Reduced Financial Risk for Users: By providing upfront cost estimates, users can avoid unexpected high bills from long-running or failed agent tasks.
-
Optimized Resource Allocation: The system ensures that computational resources are used efficiently by favoring models with better token efficiency for the perceived complexity of the task, leading to lower operational costs for developers deploying these agents at scale.
-
More Robust and Reliable Agents: By incorporating cost constraints into the planning loop, agents become inherently more cautious about redundant or unproductive exploration, leading to faster convergence on solutions and higher success rates for a given budget.
Abstract
The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that require a significant amount of tokens, three questions naturally arise: (1) Where do AI agents spend the tokens? (2) Which models are more token-efficient? and (3) Can agents predict their token usage before task execution? In this paper, we present the first systematic study of token consumption patterns in agentic coding tasks. We analyze trajectories from eight frontier LLMs on SWE-bench Verified and evaluate models' ability to predict their own token costs before task execution. We find that: (1) agentic tasks are uniquely expensive, consuming 1000x more tokens than code reasoning and code chat, with input tokens rather than output tokens driving the overall cost; (2) token usage is highly variable and inherently stochastic: runs on the same task can differ by up to 30x in total tokens, and higher token usage does not translate into higher accuracy; instead, accuracy often peaks at intermediate cost and saturates at higher costs; (3) models vary substantially in token efficiency: on the same tasks, Kimi-K2 and Claude-Sonnet-4.5, on average, consume over 1.5 million more tokens than GPT-5; (4) task difficulty rated by human experts only weakly aligns with actual token costs, revealing a fundamental gap between human-perceived complexity and the computational effort agents actually expend; and (5) frontier models fail to accurately predict their own token usage (with weak-to-moderate correlations, up to 0.39) and systematically underestimate real token costs. Our study offers new insights into the economics of AI agents and can inspire future research in this direction.
Sources
- OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
- SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
- Inverse Scaling in Test-Time Compute
- Language Models (Mostly) Know What They Know
- Budget-Aware Tool-Use Enables Effective Agent Scaling
- RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
- AgentBench: Evaluating LLMs as Agents
- Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Efficient Agents: Building Effective Agents While Reducing Cost
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering