PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs

arXiv:2607.08269 · cs.AI · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs".

Jane: The paper was written by Ying Liu, Yi Ye, Quanyu Feng, Mingxi Ye, Mingtao Zhang et al. from Hong Kong Polytechnic University, Hong Kong SAR, China.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: We’ve established the title and the scale of this project; now let's look at what the summary says PolyUQuest actually does in terms of its core mechanism. It seems to be building this massive heterogeneous graph structure over the entire website, which is a big leap from just using chunks of text.

Jane: Think about it this way; instead of just taking all the words on a page and putting them in a bag, we are mapping every piece of content—every block—to its specific location and linking it to related entities within that structure.

Lu: I think this is the foundational work that allows us to see the web as a living, interconnected organism rather than just static documents, because we are explicitly modeling all three layers: hyperlinks, page hierarchy, and entity connections.

Meng: The core idea is that by modeling the structure, we can then use a router to decide how to query this giant graph instead of dumping everything into one massive context window which would be computationally expensive.

Lalam: Lalam is excited about how the summary emphasizes that by targeting the structure, it reduces unnecessary complexity and speeds up the way AI finds what's important for cultural knowledge retrieval.

Tom: That brings us to this idea of targeted retrieval, which is where PolyUQuest really shines according to the summary; we aren't just blindly searching anymore, we are using a sophisticated map of relationships.

Jane: It’s fascinating how they unify the hyperlink topology—how pages point to each other—with the DOM hierarchy inside each page to create this one cohesive structure.

Lu: This is where I see the power; if we can follow an entity from its source block, through the page it resides on, and across links to neighboring pages, we are achieving true context.

Meng: From a practical standpoint, this means that when you’ know exactly where the information is structurally located, you avoid having to process vast amounts of irrelevant data. It's efficient data retrieval.

Lalam: It allows AI to find answers that span multiple contexts at once, making the system highly useful for people who need to understand complex topics like academic programs or research collaborations.

Tom: And this structure provides a level of clarity that simply isn't available in standard RAG setups, setting up our next discussion on how this targeted retrieval actually works in practice.

Improvements: Tom: The core improvement seems to be that the system has a two-tier router that kicks in based on how complex the question is, and this is what I think makes it so much better than older approaches by routing queries appropriately.

Jane: It’s not just one way of searching; it's different retrieval modes—Mode A for quick facts, Mode B for comparison across pages, and Mode C for deep entity reasoning—that are tailored to the specific query need.

Lu: This is a massive improvement because it means the system isn't forcing a single answer mechanism onto three distinct types of problems; we’re matching the tool to the job at hand.

Meng: And from an engineering standpoint, this routing minimizes resource use; we aren't wasting LLM tokens on questions that only need a simple lookup, which is a huge practical win for cost management.

Lalam: Lalam sees this targeted approach as enabling better cultural literacy because it allows AI to answer complex organizational questions—like finding specific professors who teach certain subjects—that require multi-step logic.

Tom: It really does, because we are not just getting an answer; we are getting verifiable provenance, meaning every citation has its source page and heading path attached so we know exactly where the evidence comes from.

Jane: Mode A is for single-hop facts, which is simple lookup, while Mode B handles those cross-page comparison queries that require connecting different parts of the site.

Lu: I think Mode C is the most impressive because it allows us to trace knowledge through entities and topics, essentially following a path of relationships rather than just retrieving text.

Meng: The two-tier router is smart because it uses lightweight rules for common patterns, only engaging the more complex LLM classifier when necessary, which keeps latency low.

Lalam: This routing ensures that we are not only getting an answer but a justification for it, providing a level of transparency that supports the educational mission of institutions.

Tom: And this blend of smart routing and targeted retrieval is what sets up the next big question: how does this approach measure up against other cutting-edge RAG systems?

Results: Tom: So, as we wrap up our discussion on PolyUQuest, the general consensus is that it solves several key issues in modern AI by being both highly structured and very efficient.

Jane: It’s not just more accurate; it' more trustworthy because we can see exactly where the data comes from and verify those claims through the citations provided.

Lu: I feel like this is a major step forward toward an AI that truly understands context and structure rather than just statistical word patterns, allowing for much deeper reasoning.

Meng: The performance gains in Table one are impressive, showing a significant jump in both correctness and faithfulness while keeping costs down—that's what matters for real-world implementation.

Lalam: Lalam believes that the ultimate impact of PolyUQuest will be seen in its deployment as a student-facing QA service at PolyU, improving how information is shared within an academic community.

Tom: And it's clear that by using this verifiable structure, we are making RAG more robust and efficient across all the different ways people use web data.

Jane: The ablation study in Table two really reinforces this, showing how critical those DOM blocks are to maintaining a coherent, structured view of the information.

Lu: It proves that without structure-aware segmentation, the system falls apart; we can't just treat the content as a flat blob of text.

Meng: The token efficiency is particularly impressive; seeing only two thousand nine hundred sixty-eight tokens per query means this is scalable and practical, unlike some other systems that inflate costs.

Lalam: This reliability allows AI to serve students better by providing answers that are grounded in verifiable truth, not just probabilistic guesses.

Tom: It’s clear that the combination of structure-aware retrieval and verifiable provenance is what makes PolyUQuest such a strong contender in the entire field of RAG.

Conclusion: Tom: To wrap up our discussion on "PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs," it really feels like we’ve seen a major step forward in how we can build truly robust and verifiable structure-aware systems.

Jane: It’s amazing how much better this makes the whole process feel—it's not just about retrieving text snippets anymore, but understanding the underlying relationships between those pieces of information.

Lu: Exactly, because when you bring in heterogeneous graphs and structure awareness, you’re moving beyond simple keyword matching; you're building genuine knowledge models that can withstand scrutiny.

Meng: And from an engineering viewpoint, the verifiable aspect is crucial; it means we can actually trust what the system outputs, which is a massive win for any enterprise adoption.

Lalam: It fundamentally changes how AI interacts with documented knowledge, allowing it to become a genuinely reliable and accountable source of information across cultures.

Tom: I agree with Meng, the confidence score that comes from verifiable structure is what separates academic novelty from industrial utility.

Jane: You're right; it makes the whole process feel grounded because the AI can point to *why* it knows something, rather than just guessing based on probability.

Lu: Thinking about the implications, this moves RAG from being a simple information retrieval layer to becoming an active reasoning component within larger AI workflows.

Meng: If we can generalize this structure-aware approach, imagine powering everything from medical diagnostics to complex legal research tools with that level of certainty.

Lalam: This capability helps foster a global culture of trust in digital information, making advanced AI feel less like a black box and more like an intelligent partner.

Tom: It really shows the whole field needs to evolve past plain text retrieval if it wants to handle the complexity of the modern web.

Jane: We are so excited about this work on PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs, and I know our listeners are too.

Lu: It's definitely set a new gold standard for how verifiable knowledge grounding should operate going forward in AI systems.

Meng: I predict that any serious commercial product aiming for high reliability will need to incorporate these graph structures soon to compete effectively.

Lalam: This signals a shift towards an era of accountable and deeply contextualized AI knowledge, which is hugely positive for human progress overall.

Ying Liu, Yi Ye, Quanyu Feng, Mingxi Ye, Mingtao Zhang, Haoyang Li*, Chen Jason Zhang, Qing Li

Hong Kong Polytechnic University, Hong Kong SAR, China

cs.AI

Submitted: 2026-08-24

Updated: 2026-08-25

Comments: Accepted at CIKM 2026 Demo Track

Code: https://github.com/13-pieces-teen/PolyUQuest

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 87/100

The gist: This paper introduces PolyUQuest, a verifiable, structure-aware web Retrieval-Augmented Generation (RAG) framework designed to overcome the limitations of existing systems that treat web pages as

Key concepts

Heterogeneous Graph
Instead of treating the web as simple chunks of text, this system maps every piece of content—every block—to its specific location and links it to related entities. This structure models hyperlinks, page hierarchy, and entity connections together.
Targeted Retrieval
PolyUQuest uses a sophisticated map of relationships to find information. Instead of dumping everything into a massive context window, the the system targets relevant data based on its structural location, which significantly reduces complexity and speeds up retrieval.
Two-Tier Router
The system employs a two-tier router that determines how to query the graph based on question complexity. It uses different retrieval modes (A for facts, B for cross-page comparison, C for deep reasoning) to match the tool to the specific problem.
Verifiable Provenance
The system provides clarity by attaching verifiable evidence to every citation. This includes the source page and a path through which information is found, allowing users to know exactly where data comes from.

Terminology

Summary

This paper introduces PolyUQuest, a verifiable, structure-aware web Retrieval-Augmented Generation (RAG) framework designed to overcome the limitations of existing systems that treat web pages as flat text. By modeling websites as heterogeneous graphs that integrate hyperlink topology, DOM hierarchy, and entity-relation knowledge, PolyUQuest enables complex reasoning across multi-page structures while providing traceable evidence for every claim.

The Core Problem

Existing RAG approaches struggle with structurally complex websites because they often flatten each page into a bag of chunks, which discards vital structural and semantic signals encoded in HTML. Plain-text and chunk-based methods fail to capture the connections between pages, while agentic web search incurs prohibitive token cost and latency. Furthermore, current Graph RAG systems often inject large global contexts that inflate query costs. There is a fundamental tension between maintaining structural fidelity, cross-page reasoning, and retrieval costs, which PolyUQuest aims to resolve by unifying three distinct structural layers:

> Layer 1 (Site Graph):

Captures how webpages point to each other via hyperlinks.

> Layer 2 (Block Tree):

Organizes page content into heading-aware text blocks that preserve the local document hierarchy.

> Layer 3 (Entity Graph):

Connects extracted entities to their source blocks, related entities, and topics.

How it works

The system operates through an offline indexing phase and an online retrieval phase. During offline indexing, the system performs web crawling, constructs a DOM-based block tree by removing non-content regions (like scripts and footers), and conducts entity extraction and resolution using an LLM to ensure that mentions are normalized across the site. The online pipeline utilizes a two-tier router to classify queries based on their structural needs. This router uses lightweight rules for unambiguous signals and an LLM classifier for long-tail queries, dispatching them to one of three specialized retrieval modes:

  1. Mode A: Direct Block Retrieval for single-hop factual questions using dense and sparse search.

  2. Mode B: Navigation Retrieval for aggregation or comparison queries that require decomposing the query into subqueries and expanding to nearby linked pages.

  3. Mode C: Entity Reasoning for multi-hop queries where evidence is scattered, utilizing both entity relation expansion and topic keyword expansion.

Verifiability and Demonstration

A key contribution of PolyUQuest is its verifiable answer provenance. Unlike systems that rely on LLM-generated summaries which may not be grounded in source text, every cited block in PolyUQuest carries its source page, heading path, and entity links. This allows users to click a citation to inspect the exact page context or follow graph paths connecting people, programs, and courses. The authors demonstrate the system using 4,240 official PolyU webpages. The demonstration includes an interactive interface with a Chat view for QA and a Graph view for exploring the underlying webpage, DOM, and entity structures.

Performance and Results

Evaluated on the PolyU-Web benchmark (300 questions), PolyUQuest outperforms existing baselines including ChunkRAG, HtmlRAG, FastGraphRAG, and LightRAG. It achieves superior results in:

> Answer Correctness (Corr.):

Reaching 0.644 compared to the next-best baseline.

> Coverage (Cov.):

Achieving 0.649 for completeness of information.

> Faithfulness (Faith.):

Attaining 0.921, representing a 36-point Faithfulness gain over the next-best baseline.

Notably, PolyUQuest is highly efficient, consuming significantly fewer LLM tokens per query than its competitors—specifically 10× fewer than LightRAG—because the router selects targeted retrieval pathways rather than injecting massive global contexts. Ablation studies confirm that structure-aware segmentation via DOM blocks is the primary driver of the system's quality advantage.

Improvements for AI systems

To improve current RAG architectures based on the PolyUQuest framework, I propose transitioning from Flat-Text Chunking to a Multi-Layered Heterogeneous Graph architecture.

Here are the specific technical improvements and the resulting capabilities of the improved AI system:


  1. Improvement: Implement a Three-Layer Heterogeneous Web Graph

Instead of converting web pages into isolated text chunks, integrate three distinct data layers into a unified graph structure:

  • The Site Topology Layer (Hyperlink connectivity between pages).

  • The DOM Hierarchy Layer (Heading-aware text blocks preserving parent/child relationships).

  • The Entity-Relation Layer (Extracted entities and their semantic associations).

Improved AI System Capability Description

:---:---

The system can navigate a website like a human browser, following links from an overview page to specific sub-pages (e.g., from a Course List to Specific Admission Requirements) to find answers that are not contained on a single page.

  1. Improvement: Deploy a Two-Tier Query Router

Replace standard similarity search with a dual-stage routing mechanism:

  • Tier 1: Lightweight heuristic rules for structural patterns (e.g., compare X and Y or who is X).

  • Tier 2: LLM-based classifier to dispatch queries into one of three specialized retrieval modes: Direct Block Retrieval, Navigation Retrieval, or Entity Reasoning.

Improved AI System Capability Description

:---:---

The system drastically reduces LLM token costs and latency by avoiding Global Context Injection. It only retrieves the specific structural path required for a query rather than feeding massive amounts of irrelevant text into the prompt.

  1. Improvement: Implement Structural-Aware Retrieval Modes

Replace generic vector similarity with specialized retrieval logic:

  • Mode A (Direct): Hybrid search (Dense + BM25) on heading-aware blocks.

  • Mode B (Navigation): Query decomposition followed by graph traversal of linked pages.

  • Mode C (Entity Reasoning): Multi-hop expansion through entity relations and topic keywords.

Improved AI System Capability Description

:---:---

The system can solve complex, multi-hop reasoning tasks, such as: Which professors in the Computing department conduct NLP research AND teach related courses? This requires connecting a person entity to a research topic entity and then to a course entity.

  1. Improvement: Integrate Verifiable Provenance via Heading-Path Indexing

Enhance the metadata of every retrieved evidence block to include its source URL, its specific DOM heading path (e.g., Home > Faculty > Computing > Staff), and associated entity links.

Improved AI System Capability Description

:---:---

The system provides Traceable Evidence. Instead of just providing an answer, it allows users to click a citation to see the exact visual context (the specific block and its surrounding headings) on the original webpage, effectively eliminating black-box hallucinations.

Sources

Related papers