Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents

arXiv:2608.21027 · cs.AI · Submitted 2026-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Don't Solve, Just Compare".

Jane: LLM agents are increasingly used for long-horizon tasks requiring sequential reasoning and tool use, making runtime intervention crucial for improving reliability without retraining.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: To start, let's look at the title itself: "Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents." That really captures the essence of what they’re doing—limiting the auxiliary model to comparison rather than solving.

Jane: Exactly, Tom. The authors are Yanze Jiang and colleagues, and they’re pointing out that we need a way to give agents a useful direction when their proposed action is bad during runtime intervention, but existing methods often require another full solver or critic.

Lu: It’s interesting how they framed the two main hurdles they have to overcome: first, the intervention has to be more than just a simple warning, and second, it needs to consider the long-term consequences instead of just surface plausibility.

Meng: That distinction about long-term consequence versus surface plausibility is what makes me think about safety; we need interventions that actually make sense in the context of the whole plan, not just looking good on paper.

Lalam: If we can get this small comparison mechanism to work, it opens up a path for more reliable agents that don't rely on needing a massive correction module every time things drift off course. That feels like a big step toward real-world deployment.

The paper's summary: Tom: So, what’s the actual mechanism they propose? They introduce Comparison-Only Tiny Advisor, or COTA, which is a comparison-only framework where the tiny comparator judges whether sampled alternatives lead to better continuations than the actor’s proposal.

Jane: That means at any decision point after an agent proposes an action, they sample some alternative actions and use this tiny comparator to rank them against the original proposal. If at least R of those alternatives defeat it, they suggest intervention.

Lu: The core mathematical idea is that the comparator learns a function Cθ(st, aA, aB) which estimates whether action A leads to better continuation than action B in state st.

Meng: So the tiny model isn't trying to find the best path from scratch; it's just checking if option A beats option B right now based on what it has learned. That seems much less demanding computationally than solving the whole task again.

Lalam: It’s smart because it keeps the task-solving power with the original, larger actor while offloading the comparison judgment to this very focused little model. It's like giving a brilliant generalist a tiny expert referee for local disputes.

The paper's improvements: Tom: The paper suggests two big improvements for this approach. First, constructive intervention must provide more than a binary warning; it needs to give the actor something actionable to replan from.

Jane: They also emphasize that this comparison should reflect the long-term consequences of the current decision rather than just how plausible that single action looks right now.

Lu: The training method they use, same-prefix counterfactuals, is clever because it helps them train the comparator on relative continuation quality between actions from the same state without mixing in different historical context.

Meng: That training setup sounds rigorous; ensuring we get good pairwise supervision for local differences is tricky when you’re dealing with long sequences of decisions. I wonder how stable that learned comparison function is when the agent moves into a completely new area of the task space.

Lalam: The statistical guarantees they provide are what really sell this, showing that even with just a tiny model, we can have a high probability that the intervention gate will work correctly when things get tricky. That level of reliability is what makes me optimistic about its application in complex agent systems.

Conclusion: Tom: So, to wrap up the paper "Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents," the main point is that we can achieve constructive runtime intervention by narrowing the auxiliary model’s role from solving a task to comparing local alternatives.

Jane: It shows that this tiny comparator framework can be sufficient to guide substantially larger actors, proving that effective intervention doesn't have to rely on another full task solver or critic.

Lu: The paper demonstrates that candidate diversity in the comparison mechanism is crucial because repeated sampling alone doesn't give you a wide enough local set of alternatives for good intervention.

Meng: I think the efficiency gain mentioned, averaging "one point three eight times the end-to-end episode time of the original actor," shows this isn't just a theoretical curiosity; it has practical overhead that’s manageable in real systems.

Lalam: Ultimately, COTA achieves what they set out to do: it provides constructive runtime intervention without needing another task solver, and it performs well across all nine evaluation settings compared to the strongest actors we tested.

Tom: It’s a solid piece of work, and I think this comparison-only approach is going to be a key tool for building more reliable AI agents in the future.

National University of Singapore

cs.AI

Submitted: 2026-08-21

Updated: 2026-09-28

Importance score: 91/100

The gist: LLM agents are increasingly used for long-horizon tasks requiring sequential reasoning and tool use, making runtime intervention crucial for improving reliability without retraining.

Key concepts

Comparison-Only Mechanism
This framework restricts the auxiliary model's job solely to comparing two actions at the same state. It doesn't try to solve the overall problem or generate a full correction. Its only function is to judge which of several proposed actions is locally superior based on a specific quality metric.
Same-Prefix Counterfactuals
This is the training method used to teach the tiny comparator how to compare actions effectively. It involves creating training data by taking an action and then comparing sibling actions from the exact same starting state, ensuring the model learns relative differences between choices.
Constructive Intervention
The goal is for runtime intervention to provide more than just a simple 'stop' or 'go' signal. COTA aims for advice that gives the original agent a useful direction by showing which local alternatives are better, enabling the agent to replan intelligently instead of just receiving a binary warning.
Domination Rate
This refers to the true probability that one action is significantly better than another from a given state. The comparator model is trained to estimate this rate, allowing it to determine if an action proposal is truly poor by checking if enough alternatives defeat it.

Terminology

Summary

LLM agents are increasingly used for long-horizon tasks requiring sequential reasoning and tool use, making runtime intervention crucial for improving reliability without retraining. This paper introduces Comparison-Only Tiny Advisor (COTA), a framework designed to achieve constructive runtime intervention using only a tiny auxiliary model capable of pairwise action comparison, rather than independent task solving or correction. The work demonstrates that this narrow capability is sufficient to guide substantially larger actors, showing that effective intervention can be realized by reducing the auxiliary model's role from solving the task to comparing local alternatives.

Problem Formulation and Motivation

The core problem addressed is how to provide a useful direction for recovery when an LLM agent proposes a poor action during runtime intervention. Existing methods often rely on an expert solver or a critic that generates task-specific corrections, which incurs the cost of another capable solver or the capacity demands of a task-capable critic. COTA investigates whether constructive runtime intervention can be achieved with a tiny auxiliary model that neither independently solves the task nor generates a correction. The paper highlights two challenges: first, constructive intervention must provide more than a binary warning, offering a useful direction for replanning; and second, it should reflect the long-term consequence of the current decision rather than its surface plausibility.

COTA Framework: Comparison-Only Mechanism

COTA is a comparison-only framework where the learned advisor performs only pairwise action comparison. At each decision point, after the actor proposes an action, the framework samples a small set of executable alternatives and uses a tiny comparator to judge them relative to the proposal. The comparator's target is defined as:

Cθ(st, aA, aB) ≈ I[Qπ(st, aA) > Qπ(st, aB)] (Equation 4)

The framework rejects the actor's proposal when at least R of the K alternatives defeat it (Equation 6). When intervention is warranted, the predicted winners are ranked by the comparator and returned as non-binding advice. The original actor then replans from this advice to produce a new proposal. The comparator's only task is local comparison in Eq. 4, leaving planning, generation, and execution entirely with the original actor.

Comparator Training via Same-Prefix Counterfactuals

Training the tiny comparator requires supervision for the relative continuation quality of two actions from the same state. To avoid conflating action effects with history differences, training data is constructed using same-prefix counterfactual branches. This involves:

  1. Restoring a decision state and executing a branch-point action to get a return, denoted as Y (st, a).

  2. Using sibling actions (aA and aB) from the same state to construct pairwise supervision: Bb(st, aA, aB) = I[QbM(st, aA) > QbM(st, aB)] (Equation 8).

The comparator is fine-tuned to predict these pairwise labels. The margin parameter γe is used during training to treat small empirical return differences as ties rather than forcing the model to learn preference from weak distinctions.

Statistical Guarantees and Performance

The paper provides theoretical guarantees for the accuracy of the learned gate. Proposition 3.1 proves that with high probability, the learned estimate of the candidate-relative domination rate converges to the true rate:

E[ρbθ,K] − ρµ(st, at) ≤ ϵθ(st, at) + r log(2/δ) / 2K (Equation 12)

This shows that when the true domination rate is sufficiently far from the threshold R/K, the learned gate agrees with the oracle gate with probability at least 1 − δ. Experimental results across WebShop, ALFWorld, and τ3-Retail show that COTA improves all nine evaluation settings and outperforms compared baselines, even for substantially stronger actors like Qwen3.6-35B-A3B and DeepSeek-V4-Flash. The efficiency is modest, with COTA averaging 1.38× the end-to-end episode time of the original actor.

Candidate Diversity and Conclusion

The candidate mechanism is crucial for injecting local diversity into the actor's decision process, as repeated sampling alone yields a highly concentrated action distribution. COTA demonstrates that this candidate mechanism exposes substantially more diverse local alternatives for intervention, and actors adopt recommended actions with substantial probability. The central design is validated: narrowing the learned task to local comparison enables effective constructive intervention while leaving task-level replanning to the stronger actor, proving that effective constructive intervention can come from narrowing the auxiliary model’s role from solving to comparing. The final conclusion is that COTA achieves the best performance in all nine actor–environment settings and shows that "constructive runtime intervention need not rely on another task solver.

Improvements for AI systems

As a diligent researcher, I have analyzed the COTA (Comparison-Only Tiny Advisor) framework presented in this paper. The core innovation is reducing constructive runtime intervention from an open-ended task-solving problem to a local, pairwise comparison primitive using a tiny comparator model.

Here are the specific improvements and what the resulting AI system can achieve:


The proposed COTA framework fundamentally improves LLM agents by enabling high-fidelity, constructive runtime intervention without requiring the auxiliary model (the advisor) to possess task-solving capability. The primary improvement lies in decoupling when to intervene from what action to take, allowing a much smaller, less capable model to guide a powerful actor.

Here are the specific improvements and capabilities of the improved AI system:

The system gains the ability to perform high-quality, constructive runtime intervention on long-horizon tasks by replacing heavy task-solving critics or expert handoffs with a lightweight, comparison-only comparator (0.5B model).

This allows for the creation of reliable safety and reliability layers that are computationally cheap and do not require retraining the main agent actor. The system no longer incurs the redundancy of needing two capable models (actor + critic/solver) to produce a correction; instead, it uses one powerful actor guided by a tiny comparison primitive.

The improved agent can identify locally suboptimal decisions in real-time that are consistently outperformed by plausible alternatives, even if those alternatives are not globally optimal for the entire task. This capability is crucial for correcting trajectory drift or local errors before they become irreversible failures, improving reliability without disrupting successful paths.

The system can generate non-binding but highly relevant advice to the main actor during execution. When the intervention gate fires, it returns a ranked list of preferred alternative actions (e.g., Action A is better than Action B), allowing the primary LLM actor to replan from that point, effectively absorbing the correction and improving its subsequent proposal.

The system can operate with minimal online overhead (averaging 1.38× end-to-end time) while achieving performance gains comparable to or exceeding baselines that use heavy intervention mechanisms (like Asym-AC or AgentPRM), especially when compared against substantially stronger actor models (e.g., Qwen3.6 or DeepSeek-V4).

The candidate mechanism broadens the local action distribution available for intervention, ensuring that the agent is exposed to a wider variety of locally relevant choices than simple repeated sampling would provide, leading to more robust interventions and better adaptation to novel states.

In summary, this improved AI system transforms from a monolithic decision-maker into a highly reliable agent capable of self-correction through local better-versus-worse comparisons guided by a tiny model, leading to superior performance across diverse interactive environments (WebShop, ALFWorld, and τ3-Retail).

Sources

Related papers