SearchWiki: Learning to Build and Navigate Knowledge Wikis for Active Information Seeking

summary

Video file (mp4)

The gist

This paper introduces SearchWiki, a framework designed to transition retrieval-augmented generation (RAG) from "passive retrieval to active corpus navigation." By addressing the fundamental flaws of

In short

The episode discusses a paper by OpenAI introducing SearchWiki, a system that replaces flat data retrieval with a hierarchical wiki structure. Hosts explain how this allows AI agents to navigate knowledge efficiently, moving from broad summaries to specific evidence. The discussion covers the use of Reinforcement Learning and its results on benchmarks like ViDoRe v3.

Key concepts

SearchWiki
SearchWiki is a system that replaces simple data retrieval with a hierarchical, typed wiki structure. This structure allows an AI agent to traverse knowledge much like a human reading a complex report by moving through three layers: document overviews, cross-document topic pages, and source records.
Active Information Seeking
This concept refers to the AI's ability to intelligently plan its search path. Instead of passively scanning all data, the the agent actively navigates from broad summaries down to specific evidence based on what it needs.
Reinforcement Learning (RL)
The system uses RL, specifically with WikiResearcher-9B, to train a navigation policy. This turns the search process into a learned skill rather than hoping for luck in retrieval, allowing the agent to learn how to be an efficient researcher.

Terminology used across episodes

This episode discusses

The paper

SearchWiki: Learning to Build and Navigate Knowledge Wikis for Active Information Seeking · Read on arXiv

OpenAI

OpenAI

Flat retrieval-augmented generation treats a corpus as a bag of chunks, discarding document hierarchy and cross document structure. We introduce SearchWiki, a harness framework that synthesizes a corpus into a hierarchical, typed, navigable wiki and trains an agent, WikiResearcher-9B, to retrieve information through multi-turn tool use. The wiki organizes knowledge into three layers - document overviews, cross- document topic pages, and page-level source records; enabling progressive refinement of retrieval when initial lookup misses. We optimize the agent's navigation policy with on-policy reinforcement learning with a multi-component reward function balancing answer correctness, retrieval quality and trajectory efficiency. Evaluation on ViDoRe-V3 (8 domains), FinanceBench, and memory benchmarks (LoCoMo, LongMemEval, PersonaMem-v2) shows that WikiResearcher- 9B which is our RL-tuned Qwen 9B model, significantly outperforms same-size untrained baselines and exceeds or matches larger external models. SearchWiki paired with WikiResearcher-9B demonstrates that learned navigation over structured corpora is a superior alternative to flat retrieval.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SearchWiki: Learning to Build and Navigate Knowledge Wikis for Active Information Seeking".

Jane: The paper was written by OpenAI from OpenAI.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: The paper summarizes a system that replaces flat retrieval with a hierarchical, typed wiki structure called SearchWiki, which allows the agent to traverse knowledge much like a human would read a complex report.

Jane: Imagine trying to find one specific detail in a giant legal code; instead of scanning every page, you start at the chapter overview and drill down into relevant sections.

Lu: That's precisely what they are doing by creating three layers: document overviews, cross-document topic pages, and source records.

Meng: I see that structure as enabling progressive disclosure—the agent only commits to reading the tiny source files when the higher levels tell it it needs that level of detail.

Lalam: This makes sense culturally because we naturally use summaries to decide where to focus our deep attention, and SearchWiki seems to have digitized that natural human behavior for AI.

Tom: So, it’s not just storing the data; the structure itself dictates how the information is sought out efficiently. But how does this system actually learn *how* to navigate?

Improvements: Jane: The core improvement is that instead of hoping the first few chunks have the answer, we train a navigation policy using Reinforcement Learning, specifically with WikiResearcher-9B.

Tom: That sounds like a massive upgrade from just hoping for luck in retrieval; it' turns the search process into a learned skill.

Lu: The way they are optimizing this policy with DAPO and RTM masking suggests that the agent is learning not just *what* to read, but *when* to backtrack when a path fails.

Meng: From an engineering standpoint, using those specific RL techniques like Decoupled Clip and Dynamic Sampling seems designed to keep the long trajectories stable during training, which is a major hurdle in multi-turn AI.

Lalam: The ability to recover from dead ends—to backtrack—is incredibly important for me because it reflects a form of intellectual resilience we want in our digital assistants.

Tom: It sounds like this system doesn' the agent learns not only how to find the answer, but how to be an efficient, intelligent researcher. This leads us directly into the results.

Results & Benchmarks: Jane: The results show that WikiResearcher-9B performs extremely well on the ViDoRe v3 dataset, achieving a seventy-one point three five percent accuracy rate in end-to-end QA.

Tom: That's impressive for an agent trained specifically on this structured wiki approach, especially when compared to the larger models that might just be brute-forcing knowledge.

Lu: And it's not just about the training domain; they also showed great generalization on FinanceBench and various memory benchmarks.

Meng: The fact that it performs well even in out-of-distribution tasks like finance suggests that the structured navigation skills are robust, which is a huge practical win for reliability.

Lalam: It’s fascinating to see it surpassing other models on memory tasks, even though SearchWiki wasn't explicitly designed for those memory types.

Tom: It really confirms that this method of active navigation is proving to be a superior strategy compared to passive retrieval techniques we used before.

Conclusion: Tom: So, we've seen how "SearchWiki: Learning to Build and Navigate Knowledge Wikis for Active Information Seeking" moves us from a simple data dump to an intelligent, navigable knowledge base.

Jane: It's clear that the whole architecture is designed to allow the AI to intelligently plan its search path, moving from broad summaries down to specific evidence.

Lu: The future of personalized AI agents feels much more structured and manageable now that they can think about their knowledge in terms of layers and connections.

Meng: I just hope the engineering implementation remains efficient, making sure this system can scale up to handle truly massive corporate knowledge bases without slowing down dramatically.

Lalam: I feel a lot of excitement about how this allows for a more sophisticated, human-like interaction with our digital tools, fostering a deeper connection to the information itself.

Tom: It's truly a paradigm shift from passive retrieval to active navigation. We hope that this "SearchWiki: Learning to Build and Navigate Knowledge Wikis for Active Information Seeking" opens the door to many new applications in the world of AI.

Lu: Absolutely, it' provides a blueprint for active learning in knowledge management.

Meng: It gives us a much better model for real-world deployment than just throwing more parameters at the problem.

Lalam: It deserves all our excitement and attention, making things better for everyone who interacts with AI.

More episodes

← Home