Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery".
Elias: The gist The study reveals that different agents vary substantially in success and cost, and higher spending does not consistently yield better outcomes.
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we're looking at a paper called "Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery". The main thing is that these agents can spend millions of tokens trying to find a vulnerability, but they often don't produce a working proof of concept.
Elias: Exactly. It digs into what consumes that budget and why those attempts fail, which is the core problem here. This study looks at two hundred traces from four different agents on CyberGym tasks, comparing an unaided baseline against four existing efficiency methods <ref:2610.11602#pg1>.
Priya: What I'm wondering about this is how much of that token spending actually leads to success versus just wasting time and tokens on dead ends.
Nadia: That’s what it gets into. The paper found some pretty clear things about how different agents behave in terms of both cost and success rates, showing that simply spending more money doesn't always mean you get better results.
Elias: They also ranked the activities that use up the most tokens and those that are the biggest roadblocks to getting a successful result. For example, code localization and understanding, along with vulnerability reasoning and trigger design, take up sixty point four percent of token consumption and are responsible for sixty-eight point seven percent of failure weight according to page two of this paper <ref:2610.11602#pg2>.
Priya: So it seems the agents spend most of their time just trying to read and understand the code, which makes sense, but that also means if that understanding part is weak, the whole process stalls.
Nadia: Right. And they showed some examples where existing efficiency methods don't consistently save money while also boosting success. In fact, they found that only about twenty-four point four percent of matched comparisons manage to keep success high while cutting the total cost down <ref:2610.11602#pg5>.
Elias: They even showed one instance where a method improved success from eighty percent to one hundred percent for Codex by lowering its cost, but it made EnIGMA's cost go up while success dropped from fifty percent to thirty percent, which shows how sensitive these interventions can be <ref:2610.11602#pg5>.
Priya: That makes sense when you think about the underlying logic; if you tweak something in a way that doesn't fit the actual code structure, it just adds overhead without adding value.
Nadia: And they pointed out that unsuitable signals are a major limitation coded into these processes, showing up in pairs across different setups <ref:2610.11602#pg7>. This suggests that simply applying a patch isn't enough; you have to understand the signal itself.
Elias: So, when we look at the solutions they propose, like AVRI—the Agent-centric Vulnerability Reasoning Interface—they are focusing on connecting input handling to unsafe operations through source locations and analysis rules <ref:2610.11602#pg9>.
Priya: How does this interface actually help the agent when it's struggling with those massive token costs? Is it just a better way to read things?
Nadia: It's about reducing the need for the agent to repeatedly read and reconstruct evidence. They introduce something called a Bidirectional Evidence Trace, or BET, which records how input moves and what conditions could make an operation unsafe <ref:2610.11602#pg2>.
Elias: That trace keeps source-supported correspondences alongside the agent's hypotheses and open questions, which aims to cut down on that repeated reading process <ref:2610.11602#pg4>.
Priya: So this is moving beyond just giving the agent more context; it’s building a persistent memory of the failure path itself.
Nadia: Right. And on the results, AVRI actually managed to reduce total cost by eighteen percent for Codex and twenty-three point seven percent for OpenCode while keeping success rates the same or even improving recall <ref:2610.11602#pg5>.
Elias: For Codex, they lowered the total cost from a baseline of one hundred thirty-two point six five down to one hundred eight point seven five, and recall went up from sixty-five percent to seventy-five percent <ref:2610.11602#pg12>. That’s a significant reduction in spending for a similar level of effectiveness.
Priya: I'm interested in the OpenCode results because that agent had a much lower baseline cost to start with, and they managed to maintain one hundred percent success while dropping the cost from four point three eight down to three point three four <ref:2610.11602#pg12>. That’s a really clean win for efficiency there.
Nadia: That's the point, Priya, it shows that lower aggregate cost can coexist with higher success on the set of successful runs themselves, and auxiliary costs can actually offset savings from the main model <ref:2610.11602#pg7>.
Elias: So when we look at these findings from "Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery," it’s clear that understanding where those millions of tokens are actually going—and what makes them fail—is crucial for making any real progress.
Priya: It points toward a future where agents don't just guess; they use a more structured way to manage the evidence they gather during discovery.
Nadia: Exactly. This paper shows that focusing on code understanding and reasoning, while using tools like AVRI to structure that evidence better, gives us real cost savings without sacrificing the quality of the vulnerability discovery <ref:2610.11602#pg5>.
Elias: So we're seeing a shift from just throwing more computational power at the problem to engineering a smarter way for agents to use their time and resources during that search <ref:2610.11602#pg8>.
Priya: I think it means that the next step isn't just another bigger model, but making sure the agent is using its knowledge in a way that respects the underlying structure of the code it’s analyzing <ref:2610.11602#pg3>.
The paper's summary: Nadia: So, basically, this paper is taking those LLM agents that try to find bugs and asking where all that massive token budget actually goes and why they keep failing <ref:2610.11602#pg8>.
Elias: Right. It shows that the biggest drain on resources isn't just one thing, but a combination of things like how fast the agent reads code and how much reasoning it does with those tokens <ref:2610.11602#pg5>.
Priya: What I find really interesting is their breakdown of what causes failure weight versus what costs tokens, because usually you think the most expensive thing is also the most important part of the job <ref:2610.11602#pg5>.
Nadia: Well, they found that code localization and understanding take up a huge chunk of those tokens—sixty percent over in some cases—but vulnerability reasoning is also really heavy on the failure side, accounting for forty percent of the total failure weight <ref:2610.11602#pg5>.
Elias: That makes sense because if you can't read the code correctly, you can't reason about vulnerabilities in it, which ties directly into my question about what assumptions an agent has when it reads that code <ref:2610.11602#pg9>.
Priya: And they look at existing ways people try to make these agents more efficient, and they found that those methods don't always work together to cut costs while keeping the success rate up <ref:2610.11602#pg5>.
Nadia: They even showed an example where one method boosted success for one agent but actually made another agent’s cost go way up and their success drop, which is a big warning sign <ref:2610.11602#pg5>.
Elias: It seems like they identified these "unsuitable signals" as a major coded limitation that limits how well any efficiency method can actually perform its job <ref:2610.11602#pg7>.
Priya: So the big implication here for someone just listening is that we need to look beyond just using bigger models or more prompting, and start thinking about how the agent actually processes and remembers the information it finds <ref:2610.11602#pg3>.
Nadia: Exactly. They propose this Agent-centric Vulnerability Reasoning Interface, AVRI, which uses something called a Bidirectional Evidence Trace to keep track of all that input propagation <ref:2610.11602#pg2>.
Elias: The idea is to stop the agent from having to read everything over and over again by storing those connections and hypotheses persistently <ref:2610.11602#pg4>.
Priya: And the results they got were pretty promising, showing that this approach actually cut total cost by about eighteen percent for some of their agents while keeping the success rates stable or even improving recall <ref:2610.11602#pg5>.
Nadia: That's a solid result, because it means we can get better results without just throwing more computational power at the problem <ref:2610.11602#pg8>.
Elias: It changes how we think about agent design, moving from just optimizing the model to optimizing the entire pipeline of how it learns and reasons about code <ref:2610.11602#pg4>.
Priya: So the next thing we should look at is how they specifically tackle those token bottlenecks, because understanding why they waste tokens is what makes this paper really useful for anyone building these kinds of systems.
The paper's improvements: Tom: So, we're looking at how they actually suggest fixing these token problems in the paper <ref:2610.11602#pg4>.
Nadia: They’re talking about building this Agent-centric Vulnerability Reasoning Interface, AVRI <ref:2610.11602#pg9>.
Elias: The core idea there is connecting the way the agent handles input directly to where it finds unsafe operations using source locations and analysis rules <ref:2610.11602#pg9>.
Priya: And they're making this interface use something called a Bidirectional Evidence Trace, or BET <ref:2610.11602#pg2>.
Nadia: That trace records the whole journey of the input, showing what happens forward and backward, along with all those assumptions and open questions <ref:2610.11602#pg4>.
Elias: So it’s supposed to stop the agent from having to repeat that reading and reconstruction process over and over again <ref:2610.11602#pg4>.
Priya: That sounds like a way to save massive amounts of processing time, which is good because those token counts add up quickly on these long discovery tasks <ref:2610.11602#pg3>.
Nadia: And they suggest a shared backend for all the source queries and analysis queries, so the agent can select what it needs from one unified interface <ref:2610.11602#pg4>.
Elias: That means instead of the AI generating all those disparate search commands separately, it uses this centralized system to handle reading, inspection, analysis, and evidence reuse all at once <ref:2610.11602#pg4>.
Priya: It sounds like they're trying to tackle that token consumption issue by making the agent's memory smarter and more structured <ref:2610.11602#pg3>.
Nadia: They also want to improve how the agent handles that heavy code localization and understanding part, which they say is really the biggest token consumer <ref:2610.11602#pg5>.
Elias: So, it’s not just about adding more context; it’s about refining the boundary between general code reading and the specific vulnerability reasoning part of the task <ref:2610.11602#pg5>.
Priya: And they use those failure bottlenecks they found earlier to prioritize what kind of help an agent should give when things go wrong, focusing specifically on those unresolved trigger conditions <ref:2610.11602#pg5>.
Nadia: It’s about making the intervention more targeted based on where the failure actually happened in the code flow <ref:2610.11602#pg8>.
Elias: So, it sounds like they’re moving toward a system where the AI doesn't just guess; it uses its memory and a structured interface to build and reuse evidence instead of just searching blindly <ref:2610.11602#pg4>.
Priya: If this works as well as they claim, it changes how we think about building these agents—it’s about engineering the agent's workflow for efficiency rather than just relying on the raw model power <ref:2610.11602#pg3>.
Conclusion: Tom: So we’re wrapping up on "Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery" <ref:2610.11602#pg8>. It really shows that efficiency isn't just about getting a better model; it’s about engineering how the AI uses its time and resources during a task.
Nadia: Exactly. The whole point is that we need to understand where those millions of tokens are actually going, because if you don't know that, you can't make the AI cheaper or more reliable <ref:2610.11602#pg5>.
Elias: It changes how we think about agent design, shifting the focus from just optimizing the model to optimizing the entire workflow of how it learns and reasons about code <ref:2610.11602#pg4>.
Priya: I think it means that for anyone building these kinds of discovery tools, you have to look at the evidence tracking as a core part of the system, not just an afterthought <ref:2610.11602#pg3>.
Nadia: Right. And they proved that by using something like AVRI with that Bidirectional Evidence Trace, you can cut total cost while maintaining success rates for both Codex and OpenCode <ref:2610.11602#pg5>.
Elias: That’s a solid result because it means we can get better results without just throwing more computational power at the problem <ref:2610.11602#pg8>.
Priya: It shows that lower aggregate cost can happen alongside higher success on the set of runs that actually succeed, which is a good thing for practical application <ref:2610.11602#pg7>.
Nadia: So the main implication is that we need to move beyond just throwing more computational power at the problem and start engineering a smarter way for agents to use their time during that search <ref:2610.11602#pg4>.
Elias: It’s about making the AI's memory smarter and more structured so it doesn't have to repeat itself constantly <ref:2610.11602#pg4>.
Priya: I just think this kind of analysis is really important for privacy too, because understanding the flow of data helps you see where information might be leaking or being misused <ref:2610.11602#pg3>.
Nadia: It definitely does. So we’ve seen how token consumption drives failure bottlenecks and how a structured interface like AVRI can help mitigate those specific issues <ref:2610.11602#pg8>.
Elias: Yeah, it points toward a future where agents are designed with persistence in mind, not just for the immediate task but for long-term reasoning <ref:2610.11602#pg4>.
Priya: It’s about making sure the agent's journey through the code isn't just a black box, but something you can actually measure and improve upon <ref:2610.11602#pg3>.
Li Lu, Yanjie Zhao, Hongjie Chen, Haoyu Wang
Huazhong University of Science and Technology
cs.CR, cs.SE
Submitted: 2026-10-08
Updated: 2026-10-08
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
The gist: The gist The study reveals that different agents vary substantially in success and cost, and higher spending does not consistently yield better outcomes.
Key concepts
- Success and Cost Metrics
- These metrics measure how well an agent finds a vulnerability (success) and how much computational resource it uses (cost). The study found that simply spending more money or tokens does not guarantee better results; different agents perform differently regarding these two factors.
- Token Consumption Bottlenecks
- The process of finding vulnerabilities is heavily reliant on understanding code and reasoning. Code localization and understanding, along with vulnerability reasoning, consume the largest share of tokens and are identified as the primary points where failed runs occur due to high token usage.
- Bidirectional Evidence Trace (BET)
- AVRI uses a persistent trace that records how input moves through the system. It tracks not just what happens, but also potential unsafe conditions, keeping source code references alongside the agent's hypotheses. This helps agents reuse evidence instead of re-reading everything.
- Existing Efficiency Methods Limitations
- Current methods to improve efficiency often fail because they rely on unsuitable signals or add unnecessary overhead. For example, one method improved success but increased cost for one agent while severely damaging the other's performance, showing that a one-size-fits-all approach doesn't work.
Terminology
Summary
The gist The study reveals that different agents vary substantially in success and cost, and higher spending does not consistently yield better outcomes.
Empirical Study Design
The research investigates three research questions: RQ1 evaluates agents’ effectiveness and cost through success, recall, elapsed time, and billed cost; RQ2 examines token consumption across activity stages and the bottlenecks observed in unsuccessful runs; RQ3 assesses the benefits of existing efficiency methods and the mechanisms that limit them. The study uses CyberGym [16], a benchmark of real-world vulnerabilities in open-source C/C++ projects, where each task requires an agent to construct an input that triggers a defect in a vulnerable build. Success requires at least one submitted input to produce a vulnerable-build exit status other than 0 or the timeout code 300, and recall additionally requires the same input to exit with status 0 on the fixed build.
Token Consumption and Failure Bottlenecks
Code localization and understanding, together with vulnerability reasoning and trigger design, account for 60.4% of token consumption and represent the two leading bottlenecks in failed runs. The combined share ranges from 44% in Codex to 71% in EnIGMA (Page 2). Vulnerability reasoning accounts for 17.3% of tokens and 40.3% of failure weight, while code localization and understanding account for 28.4% of failure weight (Page 5). The two leading bottlenecks correspond to the stages with the highest token consumption, but their ordering differs (Page 5).
Existing Efficiency Methods Analysis
The evaluated methods do not consistently preserve success while reducing cost. Only 24.4% of matched comparisons preserve success at lower total cost; unsuitable signals and auxiliary overhead limit the benefits of existing methods (Page 2). For instance, PAgent reduces Codex’s total cost from a baseline of 21.43 to 17.00 while increasing success from 80% to 100%, but increases EnIGMA’s cost from 21.06 to 23.22 while success falls from 50% to 30% (Page 5). Unsuitable signals are the most frequent coded limitation, occurring in pairs in various configurations (Page 7).
AVRI: Agent-centric Vulnerability Reasoning Interface
Motivated by these findings, AVRI is presented as an Agent-centric Vulnerability Reasoning Interface built around a persistent Bidirectional Evidence Trace (BET) (Page 2). The BET records how input propagates from the harness and what conditions could make a selected operation unsafe, retaining source-supported correspondences alongside the agent’s hypotheses and open questions (Page 2). This design aims to reduce repeated reading and reconstruction while preserving the evidence needed to construct a PoC (Page 2). AVRI connects input handling to unsafe-operation conditions through source locations and analysis rules via Algorithm 1 (Page 9).
Evaluation of AVRI Effectiveness
On 20 evaluation tasks, AVRI reduces total cost by 18.0% for Codex and 23.7% for OpenCode while preserving success rates and improving or maintaining recall (Page 5). For Codex, AVRI lowers total cost from an agent baseline of 132.65 to a new total of 108.75, with success remaining at 85% and recall increasing from 65% to 75% (Page 12). For OpenCode, AVRI lowers total cost from a baseline of 4.38 to a new total of 3.34 while maintaining success at 100% and recall at 90% (Page 12). AVRI is the only evaluated method that reduces cost for both agents without lowering either effectiveness metric (Page 5).
Conclusion
AVRI is motivated by the findings that code understanding and vulnerability reasoning dominate token use, while code localization and understanding account for the largest failure weight (Page 5). The interface connects source inspection, vulnerability reasoning, and context recovery to help agents build and reuse evidence rather than repeatedly retrieve it (Page 2). AVRI reduces total cost by 18.0% for Codex and 23.7% for OpenCode while preserving success rates and improving or maintaining recall (Page 5). The study concludes that lower aggregate cost can coexist with higher cost on the common successful subset, and auxiliary costs can offset main-model savings (Page 7). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4). This study focuses on the cost and failure bottlenecks of agents that search reachable project code and construct a crashing input from a harness entry point (Page 8). The findings describe aggregate outcomes on the 20 evaluation tasks, with one run per agent–configuration–task combination (Page 13). The paper presents AVRI as an approach to this problem (Page 4).
Improvements for AI systems
-
Bold headers: Agent-centric Vulnerability Reasoning Interface (AVRI) development. AVRI
connects input handling to these conditions through source locations and analysis rules,
enabling agents toreuse evidence instead of repeatedly retrieve and reconstruct it.
-
Tool enhancement: Implement a shared backend for source queries and analysis queries, allowing the agent to
selects queries
from a unified interface that handles reading, inspection, analysis, and evidence reuse. -
Evidence persistence: Integrate a Bidirectional Evidence Trace (BET) that
stores forward results, backward results, their supported connections, assumptions, and unresolved questions,
allowing agents torecover the saved guard n ≤ size − 2 lets it determine that the input must contain at least eleven bytes
without repeating analysis. -
Bottleneck mitigation: Focus on improving
code localization and understanding
(accounting for 43.1% of tokens) by refining the boundary between general code understanding (L) and vulnerability reasoning and trigger design (H). -
Failure diagnosis: Use the coded failure bottlenecks to prioritize interventions, specifically targeting
unresolved trigger conditions
as a high-priority area for agent assistance, sincevulnerability reasoning accounts for 40.3% of [failure weight].
-
Efficiency method refinement: Develop adaptive efficiency methods that avoid
unsuitable signals
and focus on resolving the specific limitations identified in RQ3, such as providing context recovery that is not merelyrepeated reading alone.
Sources
- Program Analysis Guided LLM Agent for Proof-of-Concept Generation
- Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases
- CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs