MDKeyChunker: What Does One LLM Call per Chunk Buy for Markdown Retrieval?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "MDKeyChunker: What Does One LLM Call per Chunk Buy for Markdown Retrieval?".
Jane: MDKeyChunker introduces a three-stage pipeline designed to enhance Retrieval-Augmented Generation (RAG) accuracy for Markdown documents by replacing multi-tool extraction passes with a single, structure-aware LLM call per chunk,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Bhavik Mangla’s paper, "MDKeyChunker: What Does One LLM Call per Chunk Buy for Markdown Retrieval?", really gets right to the heart of how we can optimize the retrieval part of RAG when dealing with structured text like Markdown. The title itself suggests a tangible benefit from making that single LLM call per chunk instead of doing multiple extraction passes.
Jane: It’s about moving away from those traditional, multi-step processes where you have to run separate models for every piece of metadata, which I know costs a lot in terms of time and resources as the document size grows. The authors are proposing a unified pipeline to solve that specific scaling issue.
Lu: The authors focus heavily on how this approach addresses three major failures in current RAG pipelines: chunk boundary fragmentation, high metadata extraction costs, and contextual isolation between chunks. They map out a clear path to solving each one sequentially across their three-stage pipeline.
Meng: I’m interested in how they handle the "contextual isolation" problem specifically; if you extract information from one chunk but don't link it to the rest of the document, that metadata becomes useless when retrieving something broader. The idea of a rolling key dictionary seems designed to fight that synonym proliferation issue where related concepts get lost.
Lalam: That rolling key propagation mechanism is interesting because it’s designed to maintain document-level context across chunks by tracking semantic keys and evicting old ones if the dictionary gets too big, capping it at forty entries. That sounds like a smart way to manage context without overwhelming the system with every piece of information.
Tom: It seems like they are proposing a very specific, structured approach rather than just throwing more data into an existing RAG framework; they are fundamentally changing how chunks are processed from the start to ensure they carry rich, linked context. This leads us into what exactly this pipeline actually does in practice.
Jane: Before we get into the mechanics of the stages, it’s important to understand that this isn't just about chunking; it’s about enriching and restructuring those chunks in a way that makes them contextually aware before they even hit the retrieval system. That sets a high bar for accuracy.
The paper's summary: Tom: So, MDKeyChunker’s core summary is a three-stage pipeline designed specifically for Markdown documents to enhance RAG accuracy by replacing multi-tool extraction passes with just one structure-aware LLM call per chunk, which then feeds into a rolling key dictionary and subsequent key-based restructuring. That’s the main takeaway from their introduction.
Jane: In simpler terms, they are saying that instead of running separate tools for summaries, entities, questions, and keywords on every small piece of text you chop off a document in the same way everyone else does, they do it all at once in one LLM interaction per chunk.
Lu: The first stage involves block parsing where they treat things like code blocks or tables as single units so they don't get split up, and then Stage Two uses that single call to pull out seven types of metadata, including a semantic key, while the rolling key dictionary keeps track of related topics across the whole document.
Meng: The second stage is really about that rolling key dictionary; it tracks context by either adding new keys or incrementing counts for existing ones, and if it hits a limit of forty entries, it drops the oldest one to keep things focused on the most relevant themes. That’s a practical way to manage document-level context.
Lalam: And then Stage Three takes those enriched chunks and uses a key-based restructuring algorithm with a first-fit bin-packing strategy based on that semantic key to merge chunks that are conceptually related, aiming to co-locate the content for better retrieval. That’s the final step of putting everything together.
Tom: So, what this means practically is that instead of having many small, isolated pieces of text floating around with their own separate metadata, you get larger, semantically coherent retrieval units formed by merging those related chunks based on the key they share.
Jane: That consolidation should mean when a user asks a complex question that spans several sections, the system pulls in all the relevant pieces at once because they’ve been grouped together by their shared semantic theme rather than being scattered across many independent retrievals.
The paper's improvements: Tom: The authors highlight several key improvements over existing methods, focusing on how this approach solves the previous problems of boundary fragmentation, high metadata costs, and contextual isolation. They show that structure-aware chunking prevents fragmentation by enforcing atomicity constraints on elements like tables and code blocks.
Jane: And they directly tackle the cost issue by showing that replacing multiple sequential extraction steps with a single LLM call cuts down the number of inference passes from O(n · m) to O(n), which is a significant improvement when dealing with large corpora.
Lu: They also introduce the innovation of rolling key propagation, which replaces what they suggest is hand-tuned scoring mechanisms with something that uses the LLM’s native semantic matching capabilities to link related content across chunks. This helps prevent synonym proliferation where different phrases for the same idea are treated as completely distinct topics.
Meng: From a practical impact view, the consolidation in Stage Three using bin-packing to merge chunks sharing a semantic key means the system isn't just retrieving more; it’s retrieving logically unified pieces of information, which should lead to much higher precision in the final generation step.
Lalam: The paper also shows that this process results in a nine point three percent reduction in chunk count when comparing the traditional method against their proposed approach on an eighteen-document corpus, and they report a high cross-reference rate of about eighty-nine point eight percent between chunks, which indicates successful context linking without inventing synonyms.
Tom: Those empirical results are pretty compelling; showing a reduction in the number of chunks while maintaining strong contextual links suggests this pipeline is much more efficient than what we’ve seen in previous RAG setups for Markdown documents.
Conclusion: Jane: To wrap up, MDKeyChunker proposes a unified three-stage process that moves away from fragmented chunking and costly multi-step metadata extraction by using a single LLM call per chunk with rolling keys to maintain context, followed by key-based merging for better retrieval units. This system should allow for much richer, more coherent answers when querying long Markdown documents.
Lu: I think the biggest contribution here is how they use structure and semantics together; it’s not just about getting more metadata, it’s about using that metadata to intelligently restructure the document into meaningful retrieval units before retrieval even happens.
Meng: From a deployment standpoint, the reduction in LLM calls and the way context is managed across chunks makes this approach much more scalable for handling large sets of technical documentation where inference cost needs to be controlled.
Lalam: And for me, I think the implication is that we can build RAG systems that are inherently more organized; by focusing on document-level context propagation through those rolling keys, we move closer to an AI system that understands a larger narrative structure rather than just a collection of isolated facts.
Tom: So, to summarize, this work on MDKeyChunker shows how treating Markdown structure as a first-class citizen and using single-call enrichment with key restructuring can lead to significantly more relevant retrieval results and better overall performance on complex documents. It’s definitely something worth watching as we build out next generation RAG systems.
cs.CL, cs.AI, cs.IR, cs.LG
Submitted: 2026-03-08
Updated: 2026-10-02
Comments: 32 pages. v3: new evaluation on Qasper and FreshStack (Laravel) with an analysis plan committed before results; results of v1-v2 withdrawn (Appendix F); title changed. Code, harness and results: https://github.com/bhavik-mangla/MDKeyChunker
Code: https://github.com/bhavik-mangla/MDKeyChunker
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: MDKeyChunker introduces a three-stage pipeline designed to enhance Retrieval-Augmented Generation (RAG) accuracy for Markdown documents by replacing multi-tool extraction passes with a single,
Key concepts
- Block parsing
- This is the first step where the system reads Markdown text and identifies six distinct structural elements: headers, code blocks, tables, lists, blockquotes, and paragraphs. The key is enforcing 'atomicity,' meaning complex structures like fenced code or tables are kept whole in a single chunk.
- Rolling key dictionary
- This dictionary tracks the document's overall context across all chunks. When an LLM enriches a chunk, it assigns a 'semantic key.' This key is added to the dictionary; if it already exists, its count increases. If the dictionary gets too large (over 40 entries), the oldest key is removed to keep only relevant, recent context.
- Key-based restructuring
- This final stage groups chunks that share the same semantic key together using a bin-packing strategy. Chunks with identical keys are merged into one larger retrieval unit. This process reduces the total number of chunks by up to 9.3% while ensuring related content is retrieved together.
- Single-call LLM enrichment
- Instead of running multiple tools for metadata extraction, MDKeyChunker uses one prompt to get seven pieces of information per chunk. This prompt receives the chunk text and the current rolling key dictionary as input, allowing the LLM to generate a title, summary, keywords, and a specific semantic key in one efficient step.
Terminology
Summary
MDKeyChunker introduces a three-stage pipeline designed to enhance Retrieval-Augmented Generation (RAG) accuracy for Markdown documents by replacing multi-tool extraction passes with a single, structure-aware LLM call per chunk, augmented by a rolling key dictionary and subsequent key-based content restructuring. This approach aims to solve the systematic failure modes of traditional RAG pipelines, including chunk boundary fragmentation, high metadata extraction costs, and contextual isolation.
The gist
MDKeyChunker proposes a three-stage pipeline for Markdown documents that (1) performs structure-aware chunking treating headers, code blocks, tables, and lists as atomic units; (2) enriches each chunk via a single LLM call extracting title, summary, keywords, typed entities, hypothetical questions, and a semantic key while propagating a rolling key dictionary to maintain document-level context; and (3) restructures chunks by merging those sharing the same semantic key via bin-packing.
Stage 1: Markdown Structural Splitting
The first stage focuses on preserving document structure by performing Block parsing.
This parser identifies six block types: header, code (fenced or indented), table, list, blockquote, and paragraph.
Crucially, the system enforces Atomicity constraints,
ensuring that elements like fenced code blocks (delimited by “‘ or), tables (including header rows and separator lines), list items, and blockquotes
are never split across chunks. Chunks are then grouped subject to size management thresholds: a minimum chunk size of τmin (default 100 characters) and a soft maximum τmax (default 1500 characters). Each resulting chunk is represented by its text, section title, content types, start line, and end line.
Stage 2: Single-Call LLM Enrichment with Rolling Keys
This stage addresses the Metadata extraction cost
problem by using a single-call LLM enrichment protocol that extracts seven metadata fields per chunk in one invocation.
The enrichment prompt is designed to receive the chunk text, section title, positional context, and the current rolling key dictionary. The design mandates a strict JSON output containing: title,
summary,
keywords
(5–8 domain-specific terms), entities
(Named entities with types like PERSON, ORG, LOC), questions
(2–3 natural questions), a specific subtopic defined as the semantic key (specific subtopic, 2–5 words, lowercase
), and related keys.
The core innovation here is the rolling key dictionary,
which maintains context across chunks. If a chunk returns a new key, it is inserted; if an existing key is returned, its count is incremented. The dictionary size is capped at Kmax=40 entries; when exceeded, the least-recently-seen key (minimum last chunk value) is evicted.
Stage 3: Key-Based Chunk Restructuring
The final stage implements a key-based restructuring algorithm that merges semantically related chunks globally via bin-packing.
This process groups chunks using the semantic key. The algorithm iterates through these groups, employing a first-fit bin-packing strategy subject to τmerge (default 3000 characters).
Chunks sharing the same key are merged into a single retrieval unit by concatenating their text and deduplicating metadata. A finalization
pass then assigns necessary attributes like position index, chunk ID (SHA-256 hash), and bidirectional navigation links to the restructured sequence C'. Chunks that remain without a key (orphans
) are handled by prepending section context to ci,
such as headers and neighbor summaries, if their size is below τorphan.
Evaluation and Performance
The paper evaluates the pipeline using an 18-document Markdown corpus (354 KB) against 30 queries. Config D, utilizing structural chunks with BM25 sparse retrieval, achieved Recall@5 = 1.000 and MRR = 0.911.
The full pipeline (Config C), which includes all three stages, reached Recall@5 = 0.867.
Empirical results show that key-based restructuring reduces the chunk count from 269 to 244 (a 9.3% reduction
). Furthermore, the system demonstrated a high cross-reference rate of 89.8%
between chunks, confirming that rolling key propagation successfully links related content without coining synonyms. The total computational complexity is analyzed as O(D + n · TLLM) for the entire pipeline, where n is the number of initial chunks and TLLM is the latency of a single LLM call.
Limitations and Future Work
The implementation has limitations, including its LLM dependency
for Stages 2 and 3, sensitivity to Key quality,
and its scope being limited to Markdown documents. The current design also has a single-document scope
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing MDKeyChunker, and what these improved systems will be able to do:
-
Do not suffer from
chunk boundary fragmentation
when processing complex Markdown documents (e.g., those containing tables, code blocks, or nested lists). -
Achieve significantly higher retrieval recall and precision by leveraging structure-aware chunking that treats headers, code fences, tables, and lists as atomic units rather than fixed-size text segments.
-
Drastically reduce the computational cost (LLM inference calls) of metadata extraction from multiple sequential passes to a single LLM call per chunk. This eliminates the scaling bottleneck of traditional pipelines (reducing complexity from O(n · m) to O(n)).
-
Improve retrieval accuracy by providing each retrieved chunk with rich, contextually relevant metadata: a descriptive title, a concise summary, domain-specific keywords, named entities (PERSON, ORG, LOC), and hypothetical questions.
-
Implement document-level context propagation using
rolling keys
to maintain topical coherence across sequential chunks. This prevents the LLM from generating synonyms for the same subtopic repeatedly and ensures that related concepts discussed in different sections are correctly linked in a single retrieval unit. -
Restructure retrieved chunks post-retrieval by merging semantically related chunks based on their extracted
semantic key
(subtopic) using a first-fit bin-packing strategy, subject to a size constraint. -
Achieve superior context density and relevance for the final LLM generation step by creating unified retrieval units that may span multiple original, structurally separate chunks (e.g., merging two fragments discussing
model types
even if they are separated by other discussion).
This improved AI system can perform the following specific tasks:
-
Analyze complex technical documentation or research papers (Markdown format) with 100% structural integrity, ensuring no critical elements like code snippets or data tables are split across retrieval boundaries.
-
Generate highly accurate, context-aware answers to
global
queries that require synthesizing information from multiple, logically related sections of a long document (e.g., answering a question about the relationship betweenapplication domains
andmodel types
). -
Ground its responses with metadata that is internally coherent (e.g., the summary accurately reflects the specific subtopic key assigned to that chunk).
-
Provide retrieval results where chunks are not just relevant based on vector similarity, but are also contextually grouped based on shared semantic themes, leading to a higher mean reciprocal rank (MRR) and recall across dense retrieval and sparse methods.
-
Operate at industrial scale with reduced latency by minimizing the number of required LLM calls per document chunk while maximizing the contextual information passed to the final generation model.
Abstract
Markdown carries structure a parser reads for free: headers, section paths, and block boundaries. Many RAG pipelines also spend LLM calls per chunk on generated metadata. We ask what one LLM call per chunk buys over that free structure. MDKeyChunker splits Markdown into header-led chunks without splitting any block; makes one LLM call per chunk for a title, summary, keywords, entities, questions, and a subtopic key, showing the model the keys already assigned in the document (a rolling key dictionary); and can merge same-key chunks. With qwen2.5:7b, we evaluate 79 Qasper questions over 30 papers and 73 FreshStack questions over 24 Laravel documentation files under BM25, two dense embedders, and hybrid fusion, following an analysis plan committed before results were computed. Evidence is matched only against source text, within a fixed token budget. Under hybrid retrieval, structural chunks beat 512-character windows on both datasets (Qasper +23.0 points, 95% CI [+12.8, +33.5]; FreshStack +5.1 [+1.4, +9.0]) and 256-token windows on Qasper (+12.7 [+5.3, +20.3]) but not on FreshStack (-2.6 [-6.4, +1.2]). Under the primary retrievers (hybrid, BM25), enrichment shows no planned-comparison difference from a free section-path prefix or from contextual retrieval; under hybrid retrieval the intervals exclude gains above 4.5 and 2.3 points on Qasper and 6.1 on FreshStack. Outside the planned comparisons, enrichment-style prefixes help BM25 on Qasper (exploratory) and mxbai on FreshStack (a secondary retriever). Rolling keys raise key reuse from 5.5% to 14.7%, but merging does not improve retrieval, and under BM25 on Qasper merging with rolling keys scores below merging without them (-6.0 [-13.1, -0.2]). Enrichment used about 1,000 input tokens per chunk; contextual retrieval 5,520 (Qasper) and 9,825 (FreshStack). The results of versions 1 and 2 are withdrawn.
Sources
- Seven Failure Points When Engineering a Retrieval Augmented Generation System
- Improving language models by retrieving from trillions of tokens
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Precise Zero-Shot Dense Retrieval without Relevance Labels
- Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
- Meta-Chunking: Learning Text Segmentation and Semantic Completion via Logical Perception
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering