Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents".
Jane: The paper was written by Thassilo M. Schiepanski and Nicholas Piel from Surfly BV.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone! Today we're digging into a paper with a great title: "Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents." Jane, I gotta say, that title alone tells you where the field is heading.
Jane: It really does, Tom. And I love the "Beyond Pixels" part because for years, the big players in web agents have been feeding screenshots to language models. You know, just showing the AI a picture of the webpage and hoping it can figure out where to click.
Tom: Right, and that works okay, but the paper's authors — Schiepanski and Piel from Surfly — they're saying, hold on, let's go back to the actual structure of the page, the DOM, the HTML underneath. The problem is, that HTML is huge.
Jane: Huge is an understatement. We're talking about files that routinely blow past the context window of even the most expensive frontier models. The paper says forty-two percent of the raw DOM snapshots they tested literally did not fit into GPT-4o's context window.
Tom: So they're stuck between a rock and a hard place, right? Screenshots are small but the model can't see them precisely enough to click the right button. The DOM has all the detail but it's too big to even process.
Jane: Exactly. And that's the puzzle this paper tries to solve. They're not saying pixels are useless, they're saying the DOM is a compelling alternative that we've been ignoring because of the size problem.
Tom: And the solution they propose is called D2Snap. It's a way to shrink the DOM down while keeping the parts that matter for an agent trying to actually do something on the page.
Jane: I think the most exciting implication here is that we might be able to move away from this hacky approach of drawing colored boxes on screenshots and hoping the vision model understands them.
Tom: You mean the Set-of-Mark stuff, the grounding approach. Yeah, that's clever but it's a workaround. If we can just feed the model clean, structured HTML, it might understand the page in a much deeper way.
Jane: And the authors show that the text alone, without the image, gets you most of the way there. That's a big deal for cost and speed, Tom.
Tom: So the title is really a manifesto. It's saying, let's stop treating the web as a visual medium for AI and start treating it as what it actually is: a structured document.
Jane: Now, the real question is, how do they actually shrink that DOM without losing the information the agent needs? That's what we're going to dig into next.
Summary: Tom: So we've set the stage. The paper is "Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents," and the core problem is that raw HTML is just too big for LLMs. Jane, walk us through what they actually built.
Jane: So they built D2Snap, which is basically a smart way to compress the DOM. It works on three levels: elements, attributes, and text. And it treats them differently because they serve different purposes.
Tom: Right, and the clever part is the element handling. They don't just randomly delete stuff. They merge nested elements by replacing a container with its own children. It's like flattening a folder structure on your computer.
Jane: And they have a really important rule: they never merge the actionable elements. Buttons, links, inputs — those are the atoms of the UI. Those stay intact. But the divs and spans that just hold them, those can be flattened away.
Tom: And for attributes, they use a scoring system. Each attribute gets a relevance score, and if it's below a threshold, it gets dropped. So things like `class` and `style` might go, but `href` and `id` stay.
Jane: Then for text, they use a summarization algorithm called TextRank. It ranks sentences by importance and keeps the top ones. So a long paragraph becomes a shorter paragraph, but the meaning is mostly preserved.
Tom: And the whole thing is controlled by three parameters, so you can dial in how aggressive you want to be on each type. The paper calls it a fidelity-size trade-off.
Jane: Now, the results are what really got me excited. They tested this on real web tasks from a dataset called Online-Mind2Web. And their reference configuration, D2Snap with parameters zero point nine, zero point three, zero point six, achieved a seventy-three percent success rate.
Tom: And the baseline, the grounded screenshot approach, got sixty-seven percent. So the DOM snapshot actually did better, even though it's using way fewer tokens.
Jane: Way fewer is right. The DOM snapshots used only sixteen point five percent of the context window on average, compared to one hundred nine percent for the raw DOM. That's the difference between fitting and not fitting at all.
Tom: So it's not just a tie, it's a win on both dimensions. Better performance and smaller size.
Jane: And there's a fascinating detail in there. When they removed the image from the grounded screenshot baseline and just gave the model the text labels, it still got sixty-two percent. So the pixels themselves are contributing almost nothing.
Tom: That's a wild finding. It suggests that the visual information is almost redundant when you have good textual grounding.
Jane: Exactly. And that's why the title says "Beyond Pixels." The future might not be about making vision models better at clicking. It might be about making text-based representations good enough that we don't need the pixels at all.
Tom: But hold on, they also found that flattening the DOM all the way to plain Markdown hurt performance significantly. So there's a sweet spot. Let's talk about what that means for how we build these agents.
Improvements: Tom: We're back with "Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents," and I want to dig into what this paper actually improves. Jane, you mentioned the sweet spot. What's the difference between a good downsample and a bad one?
Jane: The key finding is that you need to keep the HTML structure, even if it's flattened. When they went all the way to Markdown with inlined buttons, the success rate dropped to forty-six percent. But when they kept some hierarchy, even a shallow one, they got seventy-three percent.
Tom: So the structure itself carries information. The model can infer relationships between elements better when it sees them nested, even if the nesting is shallow.
Jane: And the ablation study is really telling. They tested raising each parameter to high while keeping the others low. The biggest drop in performance came from removing too many attributes. That dropped success from seventy-one percent to sixty-seven percent.
Tom: So attributes matter most. That makes sense because attributes like `href`, `id`, `aria-label` — those are the semantic hooks that tell the model what an element does.
Jane: And the element hierarchy, the nesting, that was almost free to remove. They got the same seventy-three percent success rate while cutting the size by eighteen percent. So you can flatten aggressively as long as you keep the attributes.
Tom: Now, what does this mean for the practical world? Meng, you're the engineer here. What's your take on deploying something like this?
Meng: Honestly, Tom, the latency numbers are what catch my eye. The D2Snap snapshots cut the model's processing time by about thirty percent compared to the raw DOM. That's a huge deal for real-world agents that need to respond quickly.
Jane: And the cost, right? Fewer tokens means cheaper API calls. The paper shows the reference configuration uses about twenty-one thousand tokens per snapshot versus one hundred forty thousand for raw DOM. That's a massive cost reduction.
Meng: But I also see a practical challenge. The attribute scoring table in the paper was elicited from GPT-4o itself. That's a clever trick, but it means the scoring is model-specific. A different model might score attributes differently.
Tom: That's a fair point. But the paper acknowledges that. They call it a model-informed prior. And the framework they built for evaluation is model-agnostic, so you could re-elicit the scores for any model.
Jane: And that evaluation framework is another contribution. They built a dataset of fifty-two annotated snapshots with verified solution trajectories. That's a reusable asset for the community.
Meng: Right, and it's static, so you can test different snapshot representations without running a full agent on the live web. That's a much cleaner experimental setup.
Tom: So the improvements here are threefold: a compression algorithm, an evaluation framework, and a dataset. That's a complete package for anyone working on web agents.
Jane: And the implication is that DOM-based agents are now viable. We don't have to rely on screenshots anymore. We can go back to the structured data and get better results with less compute.
Meng: I'm still curious about one thing though. How does this handle modern web apps that are heavily JavaScript-driven? The DOM can change dynamically, and the paper's dataset is mostly static pages.
Tom: That's a great question, Meng, and it's exactly what we should explore as we wrap up.
Conclusion: Tom: Alright, we're wrapping up our discussion on "Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents." Jane, give us the final summary.
Jane: So the paper tackles the fundamental problem of feeding web pages to LLMs. Raw DOM is too big, screenshots are too imprecise. D2Snap finds a middle ground by intelligently compressing the DOM while preserving what matters: actionable elements and their attributes.
Tom: And the results speak for themselves. seventy-three percent success rate on real web tasks, beating the screenshot baseline, while using only sixteen point five percent of the context window. That's a win on both fronts.
Jane: The key insight is that structure matters. You can flatten the hierarchy aggressively, but you need to keep the attributes and you need to keep some HTML, not just plain text.
Tom: And the broader implication is that we might be moving away from the pixel-based approach entirely. The paper suggests that text-based representations, when done right, can carry all the information an agent needs.
Jane: Now, Meng raised a good point about dynamic web apps. The paper's dataset is static, so there's a question about whether this works on pages that change after user interaction.
Tom: That's definitely a direction for future work. Along with testing on more models and more websites. The authors acknowledge their evaluation is limited to GPT-4o and fifty-two snapshots.
Jane: But as a proof of concept, it's compelling. It shows that DOM downsampling is not just feasible, it's actually superior to the current state of the art.
Tom: And the code is open source, so anyone can try it. That's how good research spreads.
Jane: So we'll say goodbye to "Beyond Pixels" and get ready for the next paper. Thanks for joining us, everyone.
Tom: Until next time, keep exploring beyond the pixels.
Surfly BV
cs.AI, cs.CL, cs.HC
Submitted: 2025-08-06
Updated: 2026-08-31
Comments: 21 pages, LaTeX, print version; enhanced algorithm, settled on non-inferiority claim, added downsampling monotonicity and linearity results
Code: https://github.com/webfuse-com/D2Snap
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 56/100
Key concepts
- DOM Downsampling (D2Snap)
- D2Snap is a compression technique that shrinks the raw HTML structure. It intelligently merges nested elements while preserving actionable UI components like buttons and links. It also uses a scoring system to remove irrelevant attributes and applies TextRank to condense long paragraphs, making the data manageable for LLMs.
- LLM-Based Web Agents
- These are AI agents designed to interact with websites. The paper aims to improve their functionality by replacing large, unstructured HTML inputs with smaller, structured versions that fit within the model's context window, allowing for better decision-making and interaction.
- Grounded Screenshot Approach
- This is the traditional method where an AI agent receives a visual image of a webpage. The agent must then interpret this picture to determine where to click or how to interact with the page. The study found this baseline approach achieved 67% success, which was lower than the D2Snap method.
Terminology
Summary
Summary
This paper introduces D2Snap, an algorithm for downsampling the Document Object Model (DOM) to enable the use of DOM snapshots as input to large language model (LLM)-based web agents. The authors motivate the work by noting that while DOM snapshots (serialised as HTML) are a compelling alternative to GUI snapshots because LLMs have demonstrated HTML interpretation capabilities, their excessive input token footprint has precluded reliable deployment. The paper states: "Web agents have increasingly relied on grounded GUI snapshots – screenshots augmented with visual cues – favoured for their modest input token footprint. DOM snapshots – serialised as HTML – represent a compelling alternative that leverages previously demonstrated HTML interpretation capabilities of LLMs. Their excessive input token footprint, however, has precluded reliable deployment with web agents to date."
The proposed algorithm, D2Snap, is described as an algorithm to downsample the DOM, premised on preserving actionability and actionability-discriminating features.
It operates by locally consolidating DOM nodes such that the output's utility as a snapshot is largely preserved, trading fidelity for size, while crucially remaining a valid DOM. The algorithm introduces three parameters over the unit interval: re for elements, ra for attributes, and rt for text, where 1 corresponds to maximal downsampling.
For elements, D2Snap averages elements by merging adjacent elements, governed by the parameter re which targets a DOM height of max(1, ⌈h(DOM) · (1 − re)⌉), with merge levels distributed evenly by a Bresenham gate. Actionable elements are strictly excluded from consolidation, as they represent atoms of the DOM-encoded UI – that is, triggers for UI state transitions.
Text-formatting elements are translated to token-efficient Markdown equivalents. For attributes, D2Snap discards attributes whose relevance falls below the threshold ra, using an empirical parameter A – a lookup table or scoring function calibrated against global attribute frequency. For text nodes, D2Snap consolidates text at the sentence level using the TextRank algorithm, pruning the floored rt fraction of sentences with the lowest centrality while preserving their original order, with at least one sentence always retained.
The evaluation uses a dedicated snapshot dataset sourced from Online-Mind2Web, sampling 6 easy, 6 medium, and 6 hard web browsing tasks. A human (non-author) solved each task, collecting solution trajectories of snapshots, each a triple of independent representations: a DOM snapshot, a GUI snapshot, and a grounded GUI snapshot emulating the Set-of-Mark approach. All 18 trajectories passed verification via WebJudge. The dataset contains 52 tuples of tasks and snapshot-triples, each treated as an independent sub-task. Two software engineers independently annotated every snapshot triple, identifying coherent sets of input action targets, yielding an inter-annotator agreement of F1 = 0.89.
The experimental framework deploys a minimal web agent – merely a model with a snapshot-variant system prompt
– using GPT-4o (gpt-4o-2024-11-20) as the backend, with a context window of 128 × 103 tokens. The agent's response is counted as a success if the suggested action targets form a superset of any coherent target set in the reference. The evaluation compares the baseline grounded GUI snapshots against grounding text alone, raw GUI snapshots, raw DOM snapshots, and a range of D2Snap-downsampled DOM snapshots, including a linearised Markdown representation with inlined actionable HTML.
Key results: Raw DOM snapshots exceed the model's context window in 42% of cases, at a mean context utilisation of 109%. All D2Snap-downsampled snapshots of the reference configuration fit, at a mean context utilisation of 16.5% (against 3.0% for the grounded GUI snapshot baseline). The reference configuration D2Snap.9,.3,.6 achieves a success rate of 73%, which is not significantly different from the baseline's 67% rate (+5.8%pt, 95% CI −13.6 to +26.0%pt; McNemar, p = 0.47), excluding a deficit (one-sided 95%) beyond 11%pt. Text-linearised representations (D2Snap1.0, D2SnapMD) attain the lowest success rates among D2Snap configurations, underperforming the reference configuration significantly (42%, −30.8%pt, p < 0.001; 46%, −26.9%pt, p = 0.002). The paper notes: Our results support that HTML conveys salience: flattening the DOM to text with inlined actionable elements significantly underperforms our reference (−26.9%pt, p = 0.002).
The ablation study infers the impact of each D2Snap parameter by raising a single parameter to 0.9 from an otherwise uniform 0.3. Against the uniform D2Snap.3 reference (71%, ∆Raw = 0.24): (A) re – low retention of elements – reduces size by 18% at 73% success; (B) ra – low retention of attributes – reduces size by 53% (the largest saving) but drops to 67%, indicating attributes add most to utility; (C) rt – low retention of text – neither reduces size (1.9%) nor changes success (71%) noticeably. Notably, retaining HTML at low fidelity (D2Snap.6, 73%) outperforms full removal of HTML (D2Snap1.0, 42%; +30.8%pt, p < 0.001).
The range study records uniform configurations r = re = ra = rt ∈ 0.1,..., 1.0. Size reduction is strictly monotonic (Spearman ρ = −1.000), satisfying the strong aim, and approximates a constant reduction rate: a linear regression of mean size ratio against r over the operating range fits ∆̂Raw = 0.311 − 0.226 r (Pearson correlation −0.965, R2 = 0.93), satisfying the weak aim. The drop from 0.7 to 0.8 is due to the elimination of the high-frequency class attribute.
The paper also reports that grounding text alone (GUIgrounded image) attains a success rate of 62% (−5.8%pt; McNemar, p = 0.37), indicating that image input adds little to snapshot utility. The paper states: "We further detect no significant difference in success between grounded GUI snapshots and the grounding text alone (−5.8%pt, 95% CI −18.0 to +7.1%pt), which suggests that image input contributes little to snapshot utility."
The paper's contributions are summarised as: proposing D2Snap, an algorithm to downsample the DOM premised on preserving actionability and actionability-discriminating features; demonstrating that a minimal snapshot-variant web agent (GPT-4o) achieves a success rate not significantly different from a grounded GUI snapshot baseline at a mean context utilisation of 16.5%; showing that every D2Snap-downsampled snapshot fits within the context window while 42% of raw DOM snapshots exceed it; detecting no significant difference between grounded GUI snapshots and grounding text alone; and showing that D2Snap downsamples strictly monotonically and approximately linearly across its operating range.
Limitations acknowledged include the modest evaluation dataset size (n = 52), reliance exclusively on GPT-4o, the reference configuration being selected among results, and the limitation that DOM snapshots cannot capture cross-origin documents embedded via iframes or graphics-based applications via canvas, unlike GUI snapshots. Future work directions include automated systematic exploration of the downsampling parameter space across broader LLMs and web applications, investigating alternative attribute scoring elicitation strategies, and applying DOM downsampling beyond web agents to domains such as accessibility, content summarisation, and UI testing.
Improvements for AI systems
Based on the scientific paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
-
Implementation: Integrate a DOM preprocessing module that applies the D2Snap algorithm with configurable parameters (
re,ra,rt) to reduce DOM snapshot size before sending to the LLM. -
What it does: Reduces input token footprint by up to 90% while preserving actionability. For example, with
re=0.9, ra=0.3, rt=0.6, the system fits all snapshots within context window (mean 16.5% utilization) and achieves 73% task success rate, comparable to grounded GUI baselines. -
Implementation: Use the LLM-elicited attribute scoring table (Attachment C) as a static lookup table to discard low-relevance attributes (e.g.,
accept-charset,accesskey,coords) before serialization. -
What it does: Reduces attribute overhead by 53% without significant performance loss (67% vs 71% baseline), as attributes contribute most to snapshot utility.
-
Implementation: Replace naive text truncation with TextRank sentence ranking to retain only the most central sentences in text nodes, preserving semantic content.
-
What it does: Reduces text size while maintaining task success (71% with
rt=0.3), and shows that verbosity contributes little to utility (full text retention only improves by 3.8%pt). -
Implementation: Classify elements (e.g., buttons, links, inputs) as
ACTIONABLEand exempt them from consolidation, ensuring critical interaction targets remain intact. -
What it does: Prevents loss of clickable/typeable elements during downsampling, maintaining the agent's ability to suggest valid action targets.
-
Implementation: Convert HTML text-formatting elements (e.g., ``, `
, **`) to token-efficient Markdown equivalents.
-
What it does: Further reduces token count (e.g., from 111k to 4.7k tokens at
re=ra=1, rt=0) while retaining semantic structure, enabling use with models with smaller context windows. -
Implementation: Add a pre-flight check that estimates token count of the downsampled DOM and adjusts
re,ra,rtdynamically to stay within a target context utilization (e.g., ≤20% for headroom). -
What it does: Guarantees the snapshot fits within the model's context window (e.g., 128k tokens for GPT-4o), preventing out-of-context failures (42% of raw DOMs exceed this).
-
Implementation: Modify the agent's output schema to require shortest unique CSS selectors for DOM-based snapshots, as per the paper's evaluation framework.
-
What it does: Enables precise, reproducible action targeting without pixel coordinates, improving reliability in headless or non-visual environments.
-
Process web pages with DOMs up to 10× larger than the context window by downsampling to fit (e.g., from 111k tokens to 10.6k tokens at
D2Snap.9). -
Achieve 73% task success on real-world web browsing tasks (e.g., checking drug interactions, filling forms) with a 16.5% context utilization, leaving room for multi-step history.
-
Operate without image input (text-only snapshots) at 62% success, enabling use in low-bandwidth or accessibility-constrained environments.
-
Handle complex, nested HTML structures by flattening hierarchy while preserving actionable elements and key attributes (e.g.,
href,src,id). -
Provide consistent performance across varying page sizes (e.g., from 3.7k to 66.7k tokens at 5th–95th percentile) due to monotonic size reduction.
-
Generate valid, parseable DOM output that can be further processed (e.g., translated to accessibility trees) without syntax errors.
These improvements enable a web agent that is more token-efficient, context-safe, and robust across diverse real-world websites, while maintaining competitive task success rates compared to image-based approaches.
Sources
- Understanding HTML with Large Language Models
- Language Models can Solve Computer Tasks
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection