Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents

arXiv:2608.09944 · cs.HC, cs.AI · Submitted 2026-07-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents".

Jane: The paper was written by Santosh Patapati from Stony Brook University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Jane, I have to say, this title grabbed me immediately. "Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents." It's a direct challenge to a lot of what the field has been doing.

Jane: It really is, Tom. And for our listeners, let me unpack that a bit. So, "UI agents" are basically programs that can operate a computer interface for you, like clicking buttons and filling forms on a website. And "navigation" is just getting from point A to point B on a page.

Tom: Right, and the paper's arguing that just getting there isn't enough, especially when you're building these tools for blind or visually impaired users. The authors, Santosh Patapati from Stony Brook, are saying that if an agent clicks a button for you, you need to understand *why* it clicked it.

Jane: Exactly. Think of it like a sighted assistant helping you in a store. If they just grab items and throw them in the cart without saying anything, you'd be lost. But if they say, "I'm picking up the milk because it's on your list," you feel in control. That's the core idea here.

Tom: And that's what makes this benchmark, which they call NeXUI, so interesting. It's not just testing whether the agent completes the task. It's testing whether the agent can explain its actions in a way that a nonvisual user can follow and approve.

Jane: It's a shift from thinking about these agents as autonomous tools to thinking about them as collaborators. The paper really hammers home that the user needs oversight, especially when an action might commit something important, like a payment.

Tom: Yeah, and that's a huge deal. I mean, the whole point of assistive tech is to give people agency, not to take it away. If the agent just does everything silently, it's not really assisting; it's replacing the user's judgment.

Jane: Precisely. And that's the gap this paper is trying to fill. It's saying, "Let's build a test that actually measures whether an agent can be a good partner in this process." And that starts with the title itself, which is a pretty bold statement about what the field has been missing.

Tom: It sets the stage for a benchmark that's about communication and trust, not just raw task completion. I'm really curious to see how they actually built this thing and what they found.

Summary: Jane: So, Tom, we've talked about the philosophy behind "Navigation Alone Is Not Enough," but let's get into the actual meat of the paper. The authors built a benchmark called NeXUI with two hundred twenty-five tasks across sixteen different interfaces.

Tom: And it's not just simple stuff. We're talking about opening settings, filling out forms, recovering from validation errors, even checking whether a change actually took effect. These are the messy, real-world things we all struggle with.

Jane: Right. And the key innovation is that the agent has to explain each step it takes. The benchmark isn't just checking the final state of the page; it's checking whether the explanation is grounded in what's actually on the screen at that moment.

Tom: Grounded is the key word there, Jane. The agent can't just say, "I'm filling out your address." It has to reference the actual state of the interface. And the paper has this really clever setup where some tasks have a "confirmation boundary."

Jane: Oh, I loved that part. For example, one task is to prepare a payment for a contact. The agent is supposed to enter all the details, but the correct behavior is to *stop* before pressing the final "Pay" button. It has to ask the user for approval.

Tom: That's such a realistic test. It's not just about reaching the goal; it's about knowing when *not* to reach the goal without checking in. That's a level of social awareness we don't usually see in these benchmarks.

Jane: And they tested some pretty powerful models on it. They used Gemini two point five Flash Lite, Gemini two point five Flash, and the newer Gemini three point five Flash. And the results are honestly a bit humbling.

Tom: They really are. The best model, Gemini three point five Flash, only managed a forty-four percent success rate on the validation set. That means it failed more than half the time on tasks that are designed to be achievable.

Jane: And the explanation scores were even more telling. They were around zero point seven three out of one point zero for the best model. So even when the agent did the right thing, it wasn't explaining itself all that well. It shows that making progress in an interface and communicating that progress are two very different skills.

Tom: It's a sobering result, but it's also exciting because it means there's a clear roadmap for improvement. The benchmark is hard enough that it's actually going to be useful for driving research forward, not just for showing off what we already have.

Improvements: Tom: Jane, we've established that these models are struggling, but what does the paper actually suggest we do about it? What's the path forward?

Jane: Well, the paper is pretty clear that one of the biggest issues is grounding. The models often claim things have happened when they haven't, or they reference elements that aren't actually there. So the improvement is in how we represent the interface to the agent.

Tom: Right, and they have this idea of giving the agent multiple views of the page. They provide the rendered screenshot, the accessibility tree, and a "reader view" of the content. The idea is that the agent can cross-reference these to build a more accurate picture.

Jane: Exactly. It's like giving someone both a map and a street-level photo. The map tells you the structure, and the photo tells you what it actually looks like. The paper suggests that future work should lean into this multimodal approach more heavily.

Tom: And there's another layer to it. The paper's evaluation isn't just about success or failure. It breaks things down into safety, efficiency, and explanation quality. So they're actively measuring whether the agent takes dangerous actions without asking, or whether it takes way too many steps.

Jane: That's a huge improvement over just saying "task completed" or "task failed." It gives researchers a granular view of *where* the agent is going wrong. Is it a navigation problem? Is it a communication problem? Or is it a judgment problem?

Tom: And that's the thing, Jane. The paper is really pushing for agents that can recognize when they're out of their depth. The confirmation boundary concept is a big deal because it forces the agent to say, "I've done my part, but this final step is yours to make."

Lu: If I could jump in here, Tom. That's the part that excites me the most. We're not just building better clickers; we're building agents that understand the *consequence* of their actions. That's a step toward actual machine reasoning about user intent and risk.

Meng: And from a practical standpoint, that's also what makes it deployable. If I'm building a system that's going to handle someone's finances, I need to know it has a hard stop before it commits a transaction. That's a feature, not a bug.

Jane: That's a great point, Meng. The benchmark is essentially codifying best practices for human-AI collaboration. It's saying that a good agent doesn't just do the job; it does the job in a way that keeps the human in the loop, informed and in charge.

Tom: And that's the real improvement this paper suggests. It's not a new algorithm or a new model architecture. It's a new way of measuring what "good" means, and that's going to push the whole field in a more responsible direction.

Conclusion: Jane: Well, Tom, we've had a great time unpacking "Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents." It really reframes what we should expect from assistive technology.

Tom: Absolutely, Jane. We started with the idea that navigation alone isn't sufficient, and we ended with a concrete benchmark, NeXUI, that tests for task success, safety, efficiency, and explanation quality all at once.

Jane: And the results are a clear call to action. The best model we looked at, Gemini three point five Flash, only hit a forty-four percent success rate, which tells us there's a lot of room for these systems to grow before they're truly reliable partners.

Tom: Right. And it's not just about getting the job done. It's about the agent being able to say, "Here's what I'm doing, and here's why," and knowing when to stop and ask for permission. That's the kind of collaboration that builds trust.

Jane: It's a benchmark that puts the user's experience front and center, which is exactly what we need more of in this space. It's not just about capability; it's about accountability.

Tom: So, we're going to say goodbye to this paper and this benchmark, but we're definitely keeping an eye on the space. The groundwork here is solid, and we're excited to see what agents can do when they're trained with these principles in mind.

Jane: Agreed. It's a thoughtful piece of work that gives us a clearer way to study agents that can truly support blind and visually impaired users in our increasingly complex digital world. Thanks for joining us, everyone.

Tom: And as always, we'll be back soon with another paper to dissect. Until then, keep questioning how we build these tools. See you next time.

Santosh Patapati

Stony Brook University

cs.HC, cs.AI

Submitted: 2026-07-01

Updated: 2026-08-12

Code: https://github.com/Soontosh/NeXUI-Data

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 44/100

Terminology

Summary

Summary

Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure behind visual layout. Recent UI agents can act on such interfaces; however, for assistive agents to be truly useful, they must behave as collaborators that keep users informed and in control, rather than as tools that simply take actions on users’ behalf. Most existing benchmarks judge systems primarily by task completion, without assessing how well they explain their actions or support user oversight.

The paper introduces NeXUI, a benchmark for assistive agents that must navigate interfaces while explaining each step in clear language for nonvisual use. NeXUI pairs realistic user goals with instrumented interface states, enabling agents to reason from both visual context and structural information. Its evaluation measures safety, efficiency, and task success, while also checking whether explanations are grounded in the interface state.

The contributions of the paper are as follows:

  • Introduction of NeXUI, a benchmark for assistive UI agents that must navigate interfaces while explaining their actions in language suited to nonvisual use.

  • A dataset of 225 tasks across 16 interfaces, pairing realistic user goals with interface states that preserve both visual context and accessibility structure.

  • An evaluation setting that goes beyond just task completion, measuring task success, safety, efficiency, and whether an agent’s explanations remain grounded in the current interface state.

  • A study of current foundation models on this benchmark, showing that NeXUI remains challenging, ensuring its usefulness as a platform for future research.

The dataset contains 225 production tasks across 16 interfaces, including 1,755 captured interface states and more than 31,000 candidate targets. Most tasks unfold over several steps. The dataset also contains 43 tasks with a confirmation boundary, where successful behavior requires the agent to stop before carrying out an important final action. Tasks are drawn from accessibility demonstrations, public service prototypes, authenticated applications, and enterprise workflows. The current release contains 23 easy tasks, 29 medium tasks, 128 hard tasks, and 45 very hard tasks.

Each task contains a sequence of captured interface states. Every state preserves the rendered page as well as information drawn from the browser structure, including the accessibility tree and a reader-oriented view of the page. The benchmark identifies the actions available in each state, each with a stable reference and information about the corresponding interface element. When the agent submits an action, NeXUI checks it against the transition encoded for the current state. A supported action advances the task to the next captured state. Actions outside the encoded transition do not create an alternative path. Each task includes a reference trajectory that defines the supported path toward the goal, with an explanation recorded for each step. Conditions tied to the final state determine whether the task was completed successfully. NeXUI provides further annotations for safety and explanation quality, including safety rules that identify actions that are forbidden or require approval from the user, and explanation rubrics that describe what the agent should communicate during the task.

NeXUI evaluates each run across task success, safety, efficiency, and explanation quality. A run is counted as passed only when the task is completed without a critical safety violation. Task success is determined through conditions tied to the requested goal, which may check the current page, the value of a field, or verify that the agent stopped at the correct point. Safety is evaluated through rules defined for each task; if the agent performs an action that is forbidden or requires confirmation without first returning control, the run receives a critical safety violation. Coordinate-based clicks are also recorded because they provide less reliable grounding than actions tied to a known interface element. Efficiency is measured by comparing the number of agent steps with the length of the reference trajectory; full efficiency credit is given when the task is completed in no more steps than the reference, with longer runs receiving a lower score, and the score reduced further when the task is not completed. Explanation quality is measured at each step, considering whether the explanation aligns with the selected action, refers to information in the current state, and connects the action to the user’s goal. When structured justification is provided, NeXUI can compare claims about the expected result with the next captured state. It also measures whether the agent communicates relevant confirmation boundaries and keeps its explanation concise.

The experiments evaluate three foundation models from the Gemini Flash family: Gemini 2.5 Flash Lite, Gemini 2.5 Flash, and Gemini 3.5 Flash. The experiments use the 25 tasks in the validation split. Each model is evaluated through the text-based baseline provided with NeXUI, which does not supply the page screenshot to the model. Instead, the model receives the user goal and information about the current interface state, including available action targets and recent interaction history. The stronger input setting also provides the reader view and the ARIA representation of the current page. At each step, the model returns a structured action and a short explanation. The step limit scales with the length of the reference trajectory and is capped at 50 steps. A run ends when the model finishes the task, requests user input, submits an invalid action, or reaches the step limit.

The main results show that all three models completed every scheduled run. Gemini 3.5 Flash achieved the strongest overall performance, with a task success rate of 44.0% and a mean explanation score of 0.7359. Gemini 2.5 Flash followed closely at 40.0% success. Explanation performance followed the same overall ordering as task success, with Gemini 3.5 Flash receiving the highest explanation score, followed by Gemini 2.5 Flash. However, neither model solved a majority of the tasks. The results show that stronger explanations do not remove the broader difficulty of reliable task completion, and demonstrate that NeXUI remains challenging even for recent foundation models.

The paper concludes that current foundation models can complete some NeXUI tasks, but reliable assistive behavior remains difficult. The strongest model completed 44% of the validation tasks and also received low explanation scores. This suggests that progress in general UI interaction does not yet translate directly into systems that can communicate their behavior clearly. NeXUI makes this gap easier to study by evaluating navigation and explanation within the same interaction. Future work can extend NeXUI with more interfaces and richer transition graphs, and can study agents that use screenshots together with accessibility structure. These directions may help researchers develop systems that track interface changes more reliably while also communicating more clearly. NeXUI provides an initial foundation for this work by placing task completion, explanations, and user control within a single benchmark.

Improvements for AI systems

Based on the paper, I can implement the following specific improvements to an AI system:

Improvement: Implement a post-hoc validation layer that cross-checks every agent explanation against the actual interface state before it is shown to the user. This layer verifies that:

  • No claim about a state change is made before that change is observable

  • All referenced elements exist in the current accessibility tree or reader view

  • Explanations do not describe actions not yet taken

Resulting capability: The AI system will never tell a user I have updated your address before the update is actually reflected in the interface, preventing misleading feedback that could cause users to trust incorrect actions.

Improvement: Add a safety classifier that identifies high-stakes actions (payments, deletions, submissions, irreversible changes) and automatically inserts a pause point before executing them. The system will return control to the user with a clear summary of what is about to happen, rather than proceeding autonomously.

Improvement: Modify the input pipeline to combine three representations of the current page state:

  • The rendered visual layout (screenshot)

  • The accessibility tree (ARIA roles, labels, states)

  • A reader-oriented linear view (headings, landmarks, ordered content)

The agent will reason across all three views simultaneously, using structural information for action targeting and visual information for detecting layout-dependent changes.

Improvement: Implement a planning mechanism that compares the agent's current step count against the reference trajectory length. When the agent exceeds the expected number of steps, it will re-evaluate its strategy, prioritize actions that directly advance the goal, and avoid redundant exploration.

Improvement: Integrate a lightweight scoring function that evaluates each generated explanation in real-time against three criteria:

  • Alignment with the selected action

  • Reference to concrete elements in the current state

  • Connection to the user's stated goal

If the explanation scores below a threshold, the system regenerates it before presenting it to the user.

Improvement: Implement a pre-action safety check that blocks any action flagged as forbidden by the task's safety rules. If a forbidden action is attempted, the system will not execute it, will log the violation, and will instead generate a corrective explanation to the user.


What the improved AI system can do overall: It can navigate complex, dynamic web interfaces on behalf of blind and low-vision users while keeping them fully informed and in control. It will explain every step in clear, grounded language, pause for approval before consequential actions, avoid unsafe actions entirely, and complete tasks efficiently. This transforms it from a tool that clicks for you into a collaborative assistant that users can trust and oversee.

Sources

Related papers