Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation

arXiv:2607.04576 · cs.CL, cs.CY, cs.IR · Submitted 2026-07-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases".

Jane: LLM agents increasingly answer questions against structured knowledge bases that they themselves help maintain,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Now that we’ve talked about the setup, let’s get into the actual substance of what they found in "Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation." They tested this concept on a real seven hundred nine-page markdown wiki maintained by an LLM <ref:2607.04576#pg0>.

Jane: The summary points out that the study rigorously checked two main things: first, if the restructuring changed answer quality compared to a standard index-catalog baseline, and second, whether that cost difference depended on the agent's access condition <ref:2607.04576#pg1>.

Lu: The paper emphasizes that they kept the page bodies exactly the same across all versions using immutable git tags to make sure any measured difference was due only to how the agent reached it <ref:2607.04576#pg0>.

Meng: That content parity gate is a smart move; isolating the access structure is key when evaluating these kinds of architectural changes <ref:2607.04576#pg1>.

Lalam: It really highlights that the efficiency intuition that progressive disclosure should save money isn't automatically true because it ignores *how* the agent uses what’s available <ref:2607.04576#pg2>.

Tom: Right, so they conclude that the savings aren't just about avoiding an index load, which a capable agent can usually sidestep anyway <ref:2607.04576#pg1>.

Jane: That’s the central surprise of the study; they found that for a capable tool-using agent, the benefit really comes from more targeted access, meaning fewer pages are cited and fewer tool turns per answer <ref:2607.04576#pg2>.

Lu: The results show that under forced catalog-preload conditions, there was a fifty-eight percent cost saving compared to the baseline where no retrofit was applied <ref:2607.04576#pg2>.

Meng: That thirty-one percent reduction in cited pages and tool turns per answer is what really gives this paper its practical weight, showing how much efficiency we can gain through better access patterns <ref:2607.04576#pg2>.

Lalam: So, the summary really boils down to this: progressive disclosure isn't inherently superior in quality across the board, but it’s significantly more efficient when the agent is steered toward targeted retrieval <ref:2607.04576#pg2>.

Tom: That’s a big takeaway for us; it refutes the expectation that we can just slap a new structure on and get massive savings without understanding the agent's behavior <ref:2607.04576#pg1>.

Jane: So, while quality remained non-inferior, the way answers were generated got tighter when using gist-ranked retrieval and slug-validating tools in the self-routing regimes <ref:2607.04576#pg2>.

Lu: The study found that these specific tools produce better citation validity and less padding, even if there's a tiny bit of correctness lost <ref:2607.04576#pg2>.

The paper's summary: Tom: Moving on to what the paper suggests we should actually do with this information, the improvements focus on operationalizing these findings into actionable strategies for building these systems. They suggest several ways to structure and access these knowledge bases more intelligently <ref:2607.04576#pg1>.

Jane: The most concrete improvement is pushing for a tiered access regime where you have that compact catalog and one-line summaries, which they call progressive disclosure <ref:2607.04576#pg2>.

Lu: They explicitly recommend implementing the catalog-preload regime for agents that lack self-routing capabilities, suggesting preloading the slimmed catalog and per-page summaries can yield significant cost reductions, up to fifty-eight percent in that forced condition <ref:2607.04576#pg2>.

Meng: That’s a direct engineering suggestion; if an agent can’t figure out where to go, giving it a preloaded catalog is the way to go for cost efficiency <ref:2607.04576#pg1>.

Lalam: They also suggest integrating a keyword-ranked retrieval tool that returns ranked page gists instead of just full documents or broad vector searches, which directly addresses the targeted access mechanism <ref:2607.04576#pg2>.

Tom: So, we’re looking at combining those ideas—dynamic summary generation to help guide retrieval and a toolset that prioritizes gists over full pages <ref:2607.04576#pg1>.

Jane: And they also point toward calibrating the cost-quality trade-off by setting explicit non-inferiority thresholds, acknowledging that using gist-ranked retrieval might trade a little bit of correctness for better citation validity <ref:2607.04576#pg2>.

Lu: The paper stresses the importance of dynamic summary generation because it helps guide the agent more efficiently during retrieval tasks <ref:2607.04576#pg1>.

Meng: From a practical deployment standpoint, this means we need to build tools that actively encourage that targeted access behavior rather than relying on a monolithic index load <ref:2607.04576#pg1>.

Lalam: And the authors really stress adopting the "threat-to-validity" discipline in evaluation, meaning mandatory content parity checks to ensure we’re only measuring what we intended to measure <ref:2607.04576#pg1>.

Tom: It sounds like a clear roadmap for how to move from just having a knowledge base to actively engineering an efficient way for an AI agent to interact with it <ref:2607.04576#pg1>.

Jane: So, the key is moving away from monolithic loading and toward mechanisms that help the agent infer the best path and only pull in what’s necessary for a good answer <ref:2607.04576#pg2>.

The paper's improvements: Tom: Alright, we’re coming to the end of our discussion on "Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation," and I want to wrap up what this means for the research community. The paper confirms that when you look at cost and quality together, it’s not a simple yes or no answer <ref:2607.04576#pg1>.

Jane: It really shows that the nominal saving from avoiding an index load isn't a universal win; instead, the actual benefit comes from making more targeted access choices when an agent is operating freely <ref:2607.04576#pg2>.

Lu: The study concludes that efficiency claims are regime-specific rather than deployment-general, meaning the benefit depends heavily on how the capable agent is allowed to interact with the corpus <ref:2607.04576#pg1>.

Meng: So, for us in engineering, it means we should focus on optimizing the toolset to encourage those targeted access behaviors rather than trying to force a monolithic index load <ref:2607.04576#pg1>.

Lalam: I think this is a huge step forward because it proves that even without a massive quality drop, we can make meaningful efficiency gains by designing the information access layer smartly <ref:2607.04576#pg2>.

Tom: So, to wrap up on "Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation," the paper shows that targeted access yields roughly a thirty-one percent reduction in pages cited and fewer tool turns per answer <ref:2607.04576#pg2>.

Jane: And while quality is non-inferior overall, it gets better in self-routing regimes with tighter citation validity when using gist-ranked retrieval <ref:2607.04576#pg2>.

Lu: This finding suggests that the structure of evidence organization and claim citation alignment are separable axes that can disagree in direction depending on the context <ref:2607.04576#pg2>.

Meng: So, we're seeing a clear path toward building systems where the agent makes intelligent choices about what information to pull in rather than just blindly consuming everything <ref:2607.04576#pg1>.

Lalam: We’ve really seen how these structured approaches can translate into tangible efficiency gains, which is something that will shape the future of how we build knowledge systems <ref:2607.04576#pg1>.

Tom: That’s it for this discussion on "Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation." We’ve seen how structure and access patterns dictate the real gains in efficiency <ref:2607.04576#pg1>.

Jane: It’s been fascinating watching how these different access regimes interact with the quality of the answers <ref:2607.04576#pg1>.

Lu: We’ve got a lot more to explore, but this paper lays a solid foundation for thinking about smarter knowledge management <ref:2607.04576#pg1>.

Meng: I’m eager to see how these targeted access improvements translate into real-world performance metrics on the next iteration <ref:2607.04576#pg1>.

Lalam: It’s inspiring to see how research can pinpoint exactly where the structural improvements are most impactful for our work <ref:2607.04576#pg1>.

Conclusion: Tom: So we’ve spent some time looking at "Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation," and what I’m getting is that the real secret to efficiency isn't just loading everything upfront, but making smarter choices about how the AI agent navigates that knowledge <ref:2607.04576#pg1>.

Jane: Exactly, Tom; the core finding is that for an agent capable of figuring things out on its own, more targeted access—fewer pages cited and fewer tool turns—actually delivers better results in terms of cost and quality <ref:2607.04576#pg2>.

Lu: I’m thinking about the wild implications here; if agents can learn to be efficient navigators instead of just blind ingestors, we could see a whole new way for AI to interact with massive, messy datasets <ref:2607.04576#pg1>.

Meng: From my side in engineering, it’s about building systems that can support this targeted access; if we design the retrieval tool right to prioritize gists over full pages, that’s where we see the most practical impact <ref:2607.04576#pg1>.

Lalam: I think this finding has a huge cultural impact because it shows us that efficiency isn't always about brute force; it’s about creating an architecture that respects the agent's capabilities to find what it needs <ref:2607.04576#pg2>.

Tom: It really does, Lalam; this paper proves that we can optimize for precision over volume in these knowledge bases <ref:2607.04576#pg1>.

Jane: And the study confirms that quality remains non-inferior across the board, which is reassuring when we’re looking at these architectural shifts <ref:2607.04576#pg2>.

Lu: Even with that non-inferiority, I see a path forward where dynamic summary generation becomes standard practice because it directly aids that targeted access mechanism <ref:2607.04576#pg1>.

Meng: I’m focused on the practical side; if we can set clear quality thresholds based on these results, it gives us a measurable goal for how efficient our retrieval tools need to be <ref:2607.04576#pg2>.

Lalam: That calibration point is really important because it shows we can actually quantify the trade-off between conciseness and citation validity <ref:2607.04576#pg1>.

Tom: So, to sum up this study on "Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation," we see that the efficiency gains come from better access patterns, not just avoiding an index load <ref:2607.04576#pg1>.

Jane: That’s right, and it’s a powerful reminder that designing for intelligent navigation can lead to significant operational savings while maintaining high answer quality <ref:2607.04576#pg2>.

Lu: The future work on this will definitely need to explore how these tiered regimes integrate with the broader world of self-routing agents <ref:2607.04576#pg1>.

Meng: I'm looking forward to seeing if we can prototype that gist-ranked retrieval tool we discussed, because that’s where the tangible engineering work lies <ref:2607.04576#pg1>.

Lalam: I think this entire line of research has a huge potential to shape how AI systems are designed culturally, by emphasizing intelligent interaction over just massive data consumption <ref:2607.04576#pg2>.

Tom: Fantastic stuff, team; we’ve got some solid takeaways from "Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation." Next up on the show, we're going to look at how robustness in training policies can handle those tricky policy perturbations that affect long-horizon agents.

cs.CL, cs.CY, cs.IR

Submitted: 2026-07-06

Updated: 2026-10-07

Importance score: 90/100

The gist: LLM agents increasingly answer questions against structured knowledge bases that they themselves help maintain, and this study tests whether restructuring these knowledge bases for progressive

Key concepts

Progressive Disclosure
This is a knowledge base strategy where instead of giving an agent access to the entire 709-page wiki at once, it provides only a compact catalog and one-line summaries. The goal is to offer more detailed information only when specifically requested, aiming for efficiency.
Access Regimes
These represent different ways an LLM agent interacts with the knowledge base. They include baseline full index loading (A0), a slimmed index (A1), per-page summaries (A2), and keyword-ranked retrieval tools (A3). The study tests how these different access structures affect performance and cost.
Targeted Access
This refers to the benefit of progressive disclosure, where the agent only accesses the specific, relevant parts of the knowledge base needed for an answer. This results in fewer pages being cited and fewer tool turns required to generate a response compared to loading everything.
Non-Inferiority
In this context, it means that even though progressive disclosure might not be better overall in every single test, the quality of answers remains essentially the same when comparing the reduced version (A3) to the full version (A0). It suggests no evidence of significant overall quality degradation.

Terminology

Summary

LLM agents increasingly answer questions against structured knowledge bases that they themselves help maintain, and this study tests whether restructuring these knowledge bases for progressive disclosure changes answer quality and cost across different agent access regimes. The gist: The saving does not come from avoiding the index load, which a capable agent sidesteps anyway, but is associated with more targeted access: fewer pages cited and fewer tool turns per answer.

Study Design and Controlled Ablation

The research subjects a real 709-page markdown wiki knowledge base maintained by an LLM to test the intuition that progressive disclosure—keeping only a compact catalog and one-line summaries—should reduce costs. The study employs a preregistered ablation where four versions of the corpus differ only in how the agent reaches the content, ensuring page bodies are byte-identical across arms, frozen as immutable git tags. These arms (A0–A3) represent different access structures: A0 is the baseline with no retrofit (full index load); A1 includes a slimmed index; A2 includes per-page summaries; and A3 uses a keyword-ranked retrieval tool. The design is fully crossed, featuring four arms against three conditions: protocol-constrained agent, a free self-routing agent, and a catalog-preload regime.

Evaluation Validity and Measurement

The study prioritizes answer quality as the primary outcome while treating cost as secondary and condition-dependent. To ensure evaluation validity, the researchers employed several controls:

  1. Content parity was verified by a content-parity gate to isolate differences to access structure alone.

  2. The design, hypotheses, and analysis plan were preregistered before confirmatory runs.

  3. An LLM judge from a different family (GPT-5 grading Claude) was used, blinded to condition and arm.

  4. A human rater audited a stratified subsample, with an inter-rater agreement metric of Cohen’s κ reported as a material limitation if it fell below 0.60.

Cost Analysis by Access Condition

The study measures cost using three metrics: provider-billed dollars, logical context tokens, and a cache-robust new-token burden. The results show that A3 is significantly cheaper than the baseline A0 in every access regime:

(enforced)

(free)

(forced)

The cost saving is largest under forced catalog-preload (58%) but remains significant in enforced (30%) and free (34%) regimes, demonstrating that the saving arises from more targeted access: fewer pages cited and fewer tool turns per answer, rather than avoiding an index load.

Quality Findings and Mechanism

The primary finding is that progressive disclosure is non-inferior overall, with the A3–A0 contrast being +0.01 (95% CI −0.27, +0.26), meaning no evidence of overall quality degradation. However, this non-inferiority is condition-dependent:

(self-routing regimes)

The quality improvement is associated with better citation validity (+0.05) and less padding (+0.14), at a small correctness (−0.09) and completeness (−0.08). This suggests that gist-ranked retrieval and a slug-validating tool produce tighter, better-cited answers.

Conclusion on Efficiency

The study concludes that the nominal saving from avoiding an index load is regime-specific rather than deployment-general. The actual benefit in deployment comes from targeted access, which manifests as a 31% reduction in pages cited and a corresponding decline in tool turns per answer. This finding refutes the registered expectation of a cost-null for self-routing agents, showing they save roughly a third (30% enforced, 34% free) under progressive disclosure regimes at non-inferior or better quality. The lesson is that an efficiency claim whose stated mechanism did not survive contact with how a capable agent uses the corpus, even though the benefit itself did.

Threats and Limitations

Key threats included content-parity confound, which was controlled by construction; judge bias, mitigated by cross-family grading; and a failure of the construct validity gate, as the pooled human rater Cohen’s κ was 0.23. The study also noted that while aggregate usage showed a reduction in pages cited, it did not capture per-read telemetry, highlighting a limitation for future work focusing on direct read telemetry. The final conclusion is that non-inferiority is margin- and condition-dependent, solid under self-routing access but not established under forced catalog preload.

Improvements for AI systems

Here are specific improvements to AI systems derived from the findings in this paper:

  1. Improved Knowledge Retrieval Strategy (Targeted Access over Monolithic Index Loading): Instead of relying on agents that might blindly ingest a large monolithic index or use generic vector search, implement an access strategy based on the agent's inferred intent.

  2. Implementation of Progressive Disclosure for LLM-Maintained Wikis: Restructure knowledge bases into tiered access regimes:

  3. Catalog Preload Regime (For Agents with No Self-Routing Capability): For agents that cannot infer page locations, preloading a slimmed catalog (132 KB) and per-page summaries is beneficial, leading to significant cost reductions (up to 58% savings in the forced condition).

  4. Gist-Ranked Retrieval for Targeted Access: Integrate a keyword-ranked retrieval tool that returns ranked page gists (title, summary, lead paragraph) rather than full documents or broad vector searches. This mechanism is associated with fewer pages cited and fewer tool turns per answer.

  5. Dynamic Summary Generation: Automatically generate and store a one-line summary for every page in the knowledge base, which can be leveraged by retrieval tools to guide the agent more efficiently.

  6. Cost/Quality Trade-off Calibration: Use the measured results to set explicit non-inferiority thresholds (e.g., a margin of 0.5 on a composite quality scale) for specific use cases, acknowledging that A3 (gist-ranked retrieval) trades small amounts of correctness/completeness for better citation validity and concision.

  7. Evaluation Protocol Rigor: Adopt the threat-to-validity discipline in system evaluation, ensuring that content parity checks are mandatory during testing to isolate structural effects from content changes.

  8. Self-Routing Agent Optimization: For agents capable of self-routing (free access), focus on optimizing the toolset to encourage targeted access behavior (fewer pages cited) rather than trying to force a monolithic index load, as the latter is unnecessary for these agents.

Sources

Related papers