Agentic Test-Time Scaling for WebAgents

arXiv:2602.12276 · cs.AI, cs.CL · Submitted 2026-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Agentic Test-Time Scaling for WebAgents".

Jane: The paper was written by Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John, Surya Krishnapillai, Michael W. Mahoney et al. from University of California, Berkeley and International Computer Science Institute and Lawrence Berkeley National Laboratory.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the channel, folks! Today we're digging into a paper that's got a title that just rolls off the tongue – "Agentic Test-Time Scaling for WebAgents." Jane, I gotta say, even before we crack it open, that title tells us we're in for something meaty.

Jane: Oh, absolutely, Tom. And I love it because it's so specific. We're not talking about just any AI here. We're talking about agents – programs that actually do things in a web browser, like clicking buttons, filling out forms, navigating a site. And "test-time scaling" is the idea that you can make a model smarter at the moment it's being used, not just during its initial training.

Tom: Right, so instead of spending millions on training a bigger model, you spend a bit more compute when it's actually solving a problem. It's like the difference between hiring a chef who's been to culinary school and giving that same chef more time and ingredients to perfect a dish at the last minute.

Jane: That's a great analogy, Tom. And the paper is from a team at UC Berkeley, with folks like Nicholas Lee and Amir Gholami. They're asking a really practical question: does this "more time and ingredients" trick actually work when the task isn't a one-shot puzzle, but a long, multi-step journey through a website?

Tom: And that's the million-dollar question, isn't it? Because in a single question-answer task, you can try a few different answers and pick the most popular one. But if you're trying to buy a specific product on a shopping site, you have to make a hundred little decisions in a row. One wrong click and you're lost.

Jane: Exactly. The paper's core argument is that just throwing more compute at every single step is wasteful and often doesn't even help. It's like trying to solve a maze by running every possible path at every intersection – you'll burn a ton of energy and still might end up in a dead end.

Tom: So they're saying we need to be smart about *where* we spend that extra compute. Not on the easy steps where the agent already knows what to do, but on the tricky, high-stakes steps where it's genuinely uncertain. That's the "agentic" part – the scaling is driven by the agent's own state of mind.

Jane: And that's what makes this paper so exciting. It's not just a new trick; it's a whole philosophy for how to make agents more reliable. We're going to get into the nitty-gritty of how they figured this out, but first, let's just say this title promises a smarter, more efficient way to build web agents, and I think it delivers.

Tom: I'm already on the edge of my seat. Let's get into the summary and see how they actually proved this works.

Summary: Jane: So, Tom, we've set the stage. The paper "Agentic Test-Time Scaling for WebAgents" is all about being smart with compute. The summary they give is really clear: they found that the old-school way of scaling, which they call "uniform scaling," just doesn't cut it for these long tasks.

Tom: Right, they tested this on two benchmarks, WebArena-Lite and GoBrowse. And the results were pretty stark. If you just sample more candidate actions at every single step and take a majority vote, you hit a wall. On WebArena-Lite, going from ten candidates to twenty gave them almost nothing – like a zero point two percent improvement, even though it doubled the compute cost.

Lu: And that's the key insight, isn't it? The compute isn't the bottleneck; the decision-making is. I'm Lu, by the way, and I find this fascinating because it mirrors a problem in a lot of complex systems. You can't just brute-force your way out of a problem if your basic unit of action is flawed.

Jane: Exactly, Lu. So they realized they needed a smarter way to pick the action. Instead of just counting votes, they introduced an "Arbiter" – another LLM call that looks at all the candidate actions and the current page state, and reasons about which one is actually best.

Tom: And that helped! The Arbiter outperformed simple majority voting. But here's the twist – it wasn't always better. They found cases where the Arbiter would override a strong, correct consensus and pick a bad minority action, which would completely derail the task.

Meng: So you're telling me the fix for a dumb system is to add a smart system, but the smart system sometimes overrules the dumb system when it's actually right? That sounds like a nightmare for an engineer trying to build something reliable.

Tom: You're hitting the nail on the head, Meng. That's the central problem they had to solve. They needed a way to know when to trust the vote and when to let the Arbiter take over.

Jane: And that's where the magic comes in. They looked at the vote distribution itself. If the votes are all bunched up on one action, that's high confidence. If they're spread out across many options, that's high uncertainty. They used simple statistics like entropy and the margin between the top two choices to measure this.

Lu: It's like a weather forecast. If all the models predict sun, you trust it. If half say sun and half say rain, you need a more detailed analysis. They're using the disagreement among the "models" – the candidate actions – as a signal for when to dig deeper.

Jane: Precisely. And that leads them to their main contribution, which we'll get into next. But the summary is this: uniform scaling is inefficient, and the key to fixing it is to use the agent's own uncertainty to decide when to spend more compute.

Tom: And they've got a name for that smart allocation strategy. Let's talk about CATTS.

Improvements: Tom: Alright, so the paper "Agentic Test-Time Scaling for WebAgents" has diagnosed the problem. Now, what's the cure? It's this thing they call CATTS – Confidence-Aware Test-Time Scaling. Jane, can you break that down for us?

Jane: Sure, Tom. The idea is beautifully simple. At every step, the agent samples a bunch of candidate actions. CATTS looks at how those votes are distributed. If there's a clear winner – high confidence – it just goes with the majority vote. But if the votes are all over the place – high uncertainty – it calls in the Arbiter to make a more reasoned decision.

Meng: So it's a conditional system. A gate. If confidence is high, use the cheap path. If confidence is low, pay for the expensive, smarter path. That's a classic engineering pattern, and I love it because it's practical.

Lu: And the beauty is in the details of that gate. They didn't just use a binary "confident or not." They used the entropy of the vote distribution and the margin between the top two actions. This gives them a continuous signal, so they can tune exactly how much uncertainty triggers the Arbiter.

Tom: And the results are pretty impressive. On WebArena-Lite, they got a success rate of forty-seven point nine percent with CATTS, compared to forty-three point two percent for simple majority voting. That's a big jump. But here's the kicker – they did it while using *fewer* tokens. In one configuration, they used 405K tokens compared to 920K for majority voting.

Jane: That's the part that gets me excited. They're not just making the agent smarter; they're making it more efficient. They're concentrating the compute where it matters, on the hard, contentious steps, instead of wasting it on the easy ones.

Meng: But hold on, how sensitive is this to the threshold? If I have to tune a hyperparameter perfectly for every new website, that's a maintenance nightmare.

Jane: That's a great question, Meng. The paper actually addresses that. They ran a full sweep of thresholds, and while the best one varies, most settings still beat the baseline. They suggest a default of zero point five works well across both benchmarks. So it's not super finicky.

Lu: And that's what makes it robust. The signal they're using – vote disagreement – is a fundamental property of the model's own uncertainty. It's not a brittle feature that only works in one environment. It's a general principle: when the model is unsure, spend more time thinking.

Tom: And they even compared it to other fancy methods like DeepConf, which uses token-level probabilities. CATTS gets similar or better results but doesn't need access to those internal probabilities, which means it works with any black-box API model. That's a huge practical advantage.

Jane: So the improvement isn't just a new algorithm; it's a new way of thinking about compute allocation. It's about being a smart spender, not a big spender. And that's a philosophy that could apply far beyond just web agents.

Tom: I'm curious to see where this principle could take us next. Let's wrap this up.

Conclusion: Jane: Well, Tom, we've had a fantastic time with "Agentic Test-Time Scaling for WebAgents." Let's just recap the journey. We started with the problem that throwing more compute at web agents doesn't reliably make them better.

Tom: Right, we saw that uniform scaling hits a wall. Then we learned that using an Arbiter to reason about actions helps, but it can also overrule good decisions. And the key was to use the agent's own vote distribution as a signal for when to trust the majority and when to bring in the Arbiter.

Jane: And that's CATTS in a nutshell. It's a dynamic, confidence-aware policy that allocates compute only when it's genuinely needed. The results speak for themselves – better success rates on WebArena-Lite and GoBrowse, often with fewer tokens spent.

Lu: From a research perspective, this is a really important shift. It moves us away from "more is better" and toward "smarter is better." The idea of using the model's internal disagreement as a control signal is powerful and could be applied to many other agentic tasks, like coding or robotics.

Meng: And from an engineering standpoint, I appreciate that it's practical. It doesn't require special access to model internals, it's not overly sensitive to hyperparameters, and it gives you a clear efficiency win. That's something you could actually deploy in a product.

Lalam: I think the most impactful vision here is about accessibility and sustainability. By making test-time scaling more efficient, we lower the cost of running reliable agents. This means smaller teams and even individual developers can build powerful automation tools without needing a massive compute budget. It democratizes the ability to create sophisticated AI agents.

Tom: That's a beautiful way to put it, Lalam. So, as we say goodbye to this paper, we're not just closing a book. We're opening a door to a future where AI agents are not just powerful, but also prudent. They know when to act fast and when to think deeply.

Jane: And that's a future we can all get excited about. Thanks for joining us, everyone. We'll see you on the next episode.

Tom: Take care, folks!

Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John, Surya Krishnapillai, Michael W. Mahoney, Kurt Keutzer, Amir Gholami

University of California, Berkeley · International Computer Science Institute · Lawrence Berkeley National Laboratory

cs.AI, cs.CL

Submitted: 2026-08-14

Updated: 2026-08-18

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 63/100

Key concepts

WebAgents
These are AI programs designed to interact with websites by performing actions like clicking buttons or filling out forms. The discussion focuses on how these agents make decisions during multi-step tasks, such as navigating a shopping site.
Uniform Scaling
This is the traditional method of increasing compute power for an agent at every step. The hosts found this approach to be wasteful and ineffective for long, complex tasks because it lacks intelligence in decision-making.
Arbiter
A specific component introduced in the paper, the Arbiter is another LLM call that reviews all candidate actions and the current page state. It aims to reason about which action is best, but not always overrides a correct consensus.
CATTS (Confidence-Aware Test-Time Scaling)
This is a strategy where an agent's uncertainty determines how much compute is used. If the agent has high confidence in its actions, it uses minimal compute; if uncertain, it activates the Arbiter to make a more reasoned decision.

Terminology

Summary

Summary

This paper presents a systematic study of inference-time scaling for tool-using, long-horizon web agents, and introduces a novel technique called Confidence-Aware Test-Time Scaling (CATTS) for dynamically allocating compute in multi-step agentic tasks.

The authors begin by noting that while test-time scaling has become a standard way to improve performance for single-shot reasoning tasks, its behavior on agentic, multi-step tasks remains less well-understood. They state: small per-step errors can compound over long horizons; and we find that naive policies that uniformly increase sampling show diminishing returns.

The paper first conducts an empirical study of inference-time scaling for web agents. The experimental setup uses a ReAct-based agent with gpt-oss-120b as the base model, evaluated on WebArena-Lite (165 tasks with programmatic success checks) and GoBrowse (341 tasks evaluated with an LLM-as-a-judge). At each step, the agent samples N candidate actions, which are then semantically deduplicated into clusters, inducing a vote distribution over distinct actions.

The first key finding is that uniformly increasing per-step compute quickly saturates. The authors observe: "For both WebArena-Lite and GoBrowse, we find that simply increasing candidate-generation compute yields diminishing returns. Scaling from N =1 to N =10 improves success from 38.8% to 43.2%, but the doubling compute from N =10 to N =20 produces only 0.2% additional gain despite doubling tokens." This non-monotonic scaling was also confirmed with Plan-and-Act style agents, establishing that the issue affects different agent architectures.

The paper then investigates stronger aggregation strategies. They find that an LLM-based Arbiter, which reasons over the candidate set given the current observation, can outperform naive voting. However, they also observe that arbitration is not uniformly beneficial and that extra compute is not automatically beneficial; it matters where and how we spend it inside the loop. Specifically, they find that arbitration can overrule high-consensus decisions, leading to task failures. They quantify this: Tasks without high-consensus overrides succeed at 46.9%, compared to 35.0% for tasks with at least one such override: a significant 11.9% difference (p = 0.026, Fisher’s exact test).

The core insight of the paper is that uncertainty statistics derived from the agent's own vote distribution correlate with downstream success and provide a practical signal for dynamic compute allocation. They compute two statistics at each step: entropy (Ht) and probability margin (Δt, the gap between the top-1 and top-2 action probabilities). They find that successful trajectories tend to exhibit lower entropy and higher margins, whereas failed trajectories show the opposite trend. Furthermore, they demonstrate that arbitration provides benefit when uncertainty is high but is harmful when uncertainty is low: At low entropy (0.0–0.3), the arbiter shows a net disadvantage of −4.4%... However, at higher entropy levels, the arbiter yields positive net advantages (+4–6%).

Based on these findings, the authors introduce CATTS (Confidence-Aware Test-Time Scaling), which uses vote-derived uncertainty to allocate compute only when decisions are genuinely contentious. The method applies a threshold τ to gate arbitration: if the uncertainty score (entropy or 1−margin) is below the threshold, it uses majority voting; otherwise, it invokes the arbiter.

The results show that CATTS achieves consistent improvements while being more efficient. The authors report: CATTS with entropy gating (H; τ =0.2 for WebArena-Lite, τ =0.5 for GoBrowse) achieves 47.9% and 90.2% respectively, resulting in a 4.7% and 2.2% gain over majority vote. Notably, margin-gated CATTS achieves 47.9% success on WebArena-Lite using only 405K tokens while simultaneously reducing the number of tokens by 56% compared to majority voting (920K tokens). The paper summarizes: CATTS improves performance on WebArena-Lite and GoBrowse by up to 9.1% over React while using up to 2.3× fewer tokens than uniform scaling, providing both efficiency gains and an interpretable decision rule.

The paper also evaluates DeepConf-style confidence filtering, which uses token-level log probabilities. While DeepConf variants achieve competitive results (e.g., 43.8% on WebArena-Lite at N=10), the authors note that DeepConf requires access to token-level log probabilities, which limits applicability to API-only models where such signals are unavailable. In contrast, CATTS derives uncertainty purely from the vote distribution, making it applicable to any model that supports sampling.

The authors conclude by explaining the two regimes they identified: Regime 1: Redundancy (high consensus) where many steps are routine and additional compute produces duplicates, and Regime 2: Contention (genuine uncertainty) where multiple plausible actions compete and selection quality matters most. They summarize: "We used this insight to propose CATTS, a dynamic inference-time policy that preserves majority voting when the model is confident and invokes deeper selection only when uncertainty is high. CATTS achieves consistent improvements across configurations and benchmarks while using fewer tokens, demonstrating that vote-derived uncertainty provides a practical signal for efficient compute allocation in agentic settings."

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in an AI system, along with what the improved system can do:


Implementation: Replace the current single-sample, single-action-per-step agent loop with a multi-candidate sampling mechanism. At each step, sample N=10 candidate actions, cluster them semantically, and compute two uncertainty statistics: entropy (Ht) and top-1/top-2 margin (Δt). Then apply a gating rule: if Ht ≤ τ (e.g., 0.2) or Δt ≥ 0.7, execute the majority vote action immediately; otherwise, invoke an arbiter LLM to select the best action from the candidate set.

What the improved system can do: The agent will automatically avoid wasting compute on easy, high-consensus steps (where it would have spent tokens generating redundant candidates and running unnecessary arbitration) while concentrating reasoning effort on genuinely contentious decisions. This yields up to 9.1% higher task success and 2.3× fewer tokens consumed compared to uniform scaling.

Implementation: Insert a lightweight LLM-based deduplicator between candidate generation and vote counting. This model clusters semantically equivalent actions (e.g., N/A vs. Not found, or click login vs. click sign in) into single buckets, assigning a representative action per cluster. Use a conservative merging policy that prefers false negatives over false positives.

Implementation: Modify the arbiter prompt to include explicit instructions: "If the candidate set shows overwhelming agreement (margin > 0.7), you must select the majority action unless you have strong evidence it is incorrect. Do not override consensus without specific justification." Additionally, add a confidence threshold: if the arbiter's self-reported confidence is below 0.5, default to the majority vote.

Implementation: Implement the full CATTS policy as a reusable module: (a) sample N candidates per step, (b) deduplicate semantically, (c) compute entropy and margin, (d) if uncertainty exceeds threshold, invoke arbiter (optionally with K=5 parallel arbiter calls and majority vote among them), (e) otherwise execute majority action. Tune thresholds per benchmark (e.g., τ entropy=0.2 for WebArena-Lite, τ entropy=0.5 for GoBrowse).

Implementation: Track running averages of entropy (H̄) and margin (Δ̄) across steps. If H̄ exceeds a threshold (e.g., 0.5) or Δ̄ falls below 0.5 for two consecutive steps, flag the trajectory as at risk and switch to a more conservative policy: increase candidate count to N=20, always invoke the arbiter, and add a recovery action (e.g., go back) to the candidate set.

Implementation: Replace the fixed arbiter scaling factor K with a dynamic rule: when entropy is moderately high (0.3–0.6), use K=1 arbiter call; when entropy is very high (>0.6) or margin is very low (<0.2), use K=5–10 arbiter calls and take a majority vote among them. Cap total per-step compute at a configurable budget.

Implementation: Apply the same CATTS gating logic at the plan level for two-agent systems: sample P plan candidates, deduplicate, compute uncertainty, and either majority-vote or arbitrate before passing the plan to the executor. Use a lower threshold (τ=0.3) for plans since errors at this level propagate more severely.

The improved AI system will:

  • Adapt compute allocation in real-time based on its own uncertainty, spending tokens only where they matter

  • Achieve higher task success (up to +9.1%) while using up to 56% fewer tokens

  • Avoid catastrophic overrides of correct consensus decisions

  • Scale gracefully across benchmarks of varying difficulty (WebArena-Lite: 40–48% success; GoBrowse: 86–90%)

  • Provide interpretable decision rules (entropy/margin thresholds) that operators can tune per domain

  • Work with any LLM that supports sampling, without requiring token-level log-probability access

Sources

Related papers