How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
summary
The gist
Agentic coding tasks are uniquely expensive, consuming orders of magnitude more tokens than code reasoning and code chat tasks, and models vary substantially in token efficiency across different
In short
The study analyzed token consumption in agentic coding tasks across eight frontier LLMs to understand where agents spend money. Findings show agentic tasks are 1000x more expensive than reasoning, usage is highly variable, and accuracy peaks at intermediate costs. Agents can only predict coarse cost trends themselves, suggesting that while upfront pricing is hard, self-prediction can signal high-cost tasks for better budget alerts.
Key concepts
- Agentic Tasks
- These are complex coding workflows where an AI agent performs multi-step actions like planning, executing code, and iterating. They are uniquely expensive because agents repeatedly feed context into models from different sources, leading to a massive input/output ratio compared to simpler tasks.
- Non-Monotonic Accuracy
- This means that as you spend more tokens (increase cost), performance does not always improve. In coding, this often happens because extra computation leads to redundant edits or exploration rather than deeper reasoning, causing accuracy to peak at a moderate cost level and then plateau.
- Self-Prediction Bias
- Agents can estimate their own token usage before running a task. However, this self-prediction is biased; agents systematically underestimate both input and output tokens. While not perfectly accurate for every instance, it provides a useful coarse signal to identify which tasks are likely to be very expensive.
- Stochastic Consumption
- Token usage for the same coding problem can vary wildly across different runs, sometimes differing by 30 times. This high variance makes predicting the exact cost before execution fundamentally difficult because agent behavior becomes unstable on more complex problems.
Terminology used across episodes
This episode discusses
- How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks · Paper Radio
- OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
- SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
- Inverse Scaling in Test-Time Compute
- Language Models (Mostly) Know What They Know
- Budget-Aware Tool-Use Enables Effective Agent Scaling · Paper Radio
- RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
- AgentBench: Evaluating LLMs as Agents
- Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Efficient Agents: Building Effective Agents While Reducing Cost
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
The paper
How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks · Read on arXiv
Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea
University of Michigan · Stanford University · All Hands AI 4Google Deepmind Microsoft AI Massachusetts Institute of Technology
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks".
Tom: Agentic coding tasks are uniquely expensive, consuming orders of magnitude more tokens than code reasoning and code chat tasks, and models vary substantially in token efficiency across different frontier LLMs.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about that title again, "How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks," because it perfectly captures the tension between the hype around agent capabilities and the very real economic reality of using them.
Jane: Exactly; it frames the whole discussion around a practical problem for developers and users: if we don't know how much this coding agent will cost us, we can't truly manage our expectations or budget effectively.
Lu: The authors are pulling from some major institutions, like the University of Michigan and Stanford University, which tells you this isn't just a niche study; it’s coming from a group that understands the broader implications for how AI is deployed in complex software engineering environments.
Meng: I saw that they are analyzing trajectories from eight frontier LLMs on SWE-bench Verified, and I wonder if testing them across such a diverse set of models helps isolate behavior caused by the model architecture versus the inherent difficulty of the task itself.
Lalam: The authors are opening up all agent trajectories, which is huge because it lets everyone see exactly what these agents are doing step-by-step when they spend their resources, which is a great way to build trust.
The paper's summary: Tom: So, the main takeaway from this study on "How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks" is that agentic tasks are uniquely expensive compared to simpler things like code reasoning or code chat, consuming one thousand times more tokens <ref:2604.22750#pg0,Analyzing and Predicting Token Consumption in Agentic Coding Tasks>.
Jane: That massive difference is striking because it shows that the way agents operate—reading repositories, calling tools, and iterating—demands a huge input token ratio compared to just asking a model a question directly.
Lu: The paper also found that token usage isn't stable; runs on the same task can vary by up to thirty times in total tokens, which suggests that the process itself is inherently stochastic rather than deterministic.
Meng: That variability is something I’ve seen in testing, and it makes it really difficult for us to guarantee a fixed cost for a specific type of agentic workflow beforehand.
Lalam: And they also showed that higher token usage doesn't automatically mean better accuracy; sometimes the performance peaks at an intermediate cost point before leveling off at higher costs.
The paper's improvements: Tom: The paper proposes a few key directions for improvement, and one of the big ones is formulating a pre-execution agent token consumption prediction task to benchmark models against.
Jane: They are trying to solve that prediction problem, showing that while agents can capture coarse trends in usage, they still have a significant gap when trying to predict exact instance-level usage across different models.
Lu: They specifically highlight that output token usage is easier to predict than input token usage because the context construction adds a lot of uncertainty right at the beginning of the process.
Meng: This finding about output versus input prediction is practically useful because it tells us where we can get a more reliable signal for budget alerts, even if it’s not perfect.
Lalam: I think that ability to give us that coarse-grained signal for relative cost is a big step toward making these agents feel more manageable and predictable for the end user.
Conclusion: Tom: So, wrapping up this discussion on "How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks," we see that agentic tasks are way more token-heavy than basic reasoning, and predicting exact costs before execution is still a genuinely difficult task for the current models.
Jane: It’s clear that while self-prediction offers some useful coarse trends, the inherent stochastic nature of these workflows means we can't rely on simple upfront pricing right now.
Lu: The paper successfully illuminated where tokens go in agentic coding and what we can anticipate before execution, which is a solid foundation for future research into agent economics.
Meng: For practical implementation, the biggest lesson seems to be that consumption-based pricing will remain the most straightforward option until these prediction methods become much more robust.
Lalam: I think the ability for agents to predict their own cost at least at a coarse level is a concrete signal for providers to issue early warnings, which is something we can start building on immediately.
Tom: That’s right; it confirms that understanding token consumption patterns is essential for anyone building or using these sophisticated AI systems. We'll keep an eye on how the community develops methods to make those predictions more reliable before we move on to the next piece of research.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck