Document Optimization for Black-Box Retrieval via Reinforcement Learning

arXiv:2604.05087 · cs.CL, cs.IR · Submitted 2026-04-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Document Optimization for Black-Box Retrieval via Reinforcement Learning".

Tom: Document expansion is recast as a document optimization problem where an instruction-tuned language model or vision language model is fine-tuned to transform documents into representations that better align with the…

Jane: First, who's behind it and why it matters.

Title and authors: Tom: We’ve touched on the concept, but let's talk about the specifics of the paper's title and who came up with it. It’s "Document Optimization for Black-Box Retrieval via Reinforcement Learning." That tells us exactly what we're looking at: using reinforcement learning to optimize documents when you can only see how well a retriever ranks things, without knowing its inner workings.

Jane: And the authors are Tom Uzan, Ron Polonsky, Douwe Kiela, and Christopher Potts from Stanford University. Having researchers from such a strong group working on this really gives us confidence in the technical depth of what they've done here.

Lu: The title highlights two key ideas: document optimization and black-box retrieval. That’s a powerful combination because it tackles the problem of improving retrieval quality while acknowledging that we often don't have complete transparency into how those retrievers function internally.

Meng: So, if I understand correctly, they aren't trying to build a better retriever itself, but rather making the documents *better* for the existing one by using reinforcement learning as the training engine. That’s a different kind of engineering challenge entirely.

Lalam: The authors focusing on this approach suggests that we can achieve performance gains by manipulating the input data structure rather than just tweaking the retrieval algorithm itself, which feels like a very practical path forward for improving our AI infrastructure.

The paper's summary: Tom: Moving into what they actually did, "Document Optimization for Black-Box Retrieval via Reinforcement Learning" takes document expansion and reframes it as an optimization problem where an instruction-tuned language or vision language model fine-tunes documents to better fit the expected query distribution under a target retriever.

Jane: That’s a big concept. So, they are essentially training an AI policy to generate new document versions that we know will boost retrieval quality when those versions are scored by the target retriever. It’s about learning what the retriever "wants" from its documents.

Lu: The core of their methodology involves formulating this as a reinforcement learning problem where the policy generates candidate transformations, and these are evaluated counterfactually by measuring how much they change the ranking metrics like nDCG@k when compared to a fixed collection of queries.

Meng: The reward mechanism is what I find most intriguing, though; they define it using this counterfactual change in nDCG@k, which captures the gain on both positive and negative query sets associated with the document being optimized.

Lalam: It’s powerful because they use Group Relative Policy Optimization, or GRPO, to update the policy based on these relative reward comparisons between different rewrites without needing an explicit value function. This makes it a very efficient way to train the policy.

The paper's improvements: Tom: When we look at what they actually achieved, "Document Optimization for Black-Box Retrieval via Reinforcement Learning" suggests that their approach yields retrieval gains across single-vector, multi-vector, and lexical retrievers in both code and visual document retrieval tasks.

Jane: That’s a broad applicability. It means this technique isn't just for one type of search engine or one kind of data; it can work across different underlying architectures as long as we can check the rankings.

Lu: A key finding they highlight is that these learned document transformations don't just expand the text randomly; they actually result in documents that are more compact, with a mean of three hundred forty-six tokens compared to zero-shot expansions which average three hundred eighty-two tokens.

Meng: That compactness is significant for practical deployment because it means we aren't just bloating our data; we are selectively refining the information to be much denser and more signal-rich for the retrieval system.

Lalam: They also showed that these optimized documents tend to form clusters in the representation space, meaning they shift "closer to the query and farther from negatives," which is a very clear indication of how their policy learns to align with the target retriever's internal logic.

Conclusion: Tom: So, wrapping up this discussion on "Document Optimization for Black-Box Retrieval via Reinforcement Learning," the main implication is that we can improve retrieval quality by rewriting documents to better match what a specific retriever expects, and this works even when we only have black-box access to those rankings.

Jane: It’s a solid framework because it shows that optimizing the document space itself is a promising direction for improving search performance across various data types. We can now leverage AI to tailor our data representations specifically for our retrieval needs.

Lu: The ability to achieve these gains offline, preserving inference-time efficiency and adding no query-time complexity, is a very strong technical achievement that opens up new avenues for how we structure knowledge bases.

Meng: From a practical standpoint, this means we might be able to improve performance using less computationally expensive embedding models while still getting results comparable to larger ones. That kind of efficiency is what matters in production systems.

Lalam: I think the most exciting part is that this method can be combined with joint adaptation, where the retriever gets periodically fine-tuned on the newly transformed corpus, leading to very substantial performance increases as they showed with Jina-ColBERT-V2.

Tom: Exactly. So, we’ve seen how this paper tackles document optimization for black-box retrieval. It’s a method that shows we can refine our data representations intelligently without needing deep internal knowledge of the retriever itself.

Jane: That's a lot to digest, Tom, but it really paints a clear picture of how document structure can be tuned for better search results.

Lu: We definitely have so much more room to explore how this optimization framework can be applied beyond code and visuals into other complex domains.

Meng: It’s going to keep me busy thinking about how we can integrate this idea into our existing document processing pipelines for better signal extraction.

Lalam: It’s exciting because it suggests that optimizing the representation itself is a promising and underexplored direction in AI research right now.

Stanford University

cs.CL, cs.IR

Submitted: 2026-04-06

Updated: 2026-10-07

Importance score: 91/100

The gist: Document expansion is recast as a document optimization problem where an instruction-tuned language model or vision language model is fine-tuned to transform documents into representations that

Key concepts

Document Optimization
This is treating document rewriting as a machine learning problem. An instruction-tuned model tries different ways to rewrite a document to make it more useful for a specific search system (retriever). The goal is not just to change the text, but to change its representation so that the search engine can find it more accurately.
Counterfactual Change ($\Delta$nDCG@k)
This measures how much a document's quality changes when it is transformed. It compares the retrieval performance (measured by nDCG) of a query against the original document versus the transformed version. This calculation quantifies exactly how much better or worse the new document performs in finding relevant results.
Group Relative Policy Optimization (GRPO)
This is a reinforcement learning technique used to train the rewriting model. Instead of needing a complex value function, GRPO optimizes the policy by comparing different transformation paths relative to each other. This helps the model learn which changes are beneficial for improving retrieval quality without needing explicit feedback on every single action.
Representation Space Alignment
The policy learns to generate rewrites that position the document's meaning closer to where the retriever expects it. Analysis shows optimized documents cluster better and move closer to relevant queries while moving away from irrelevant ones, confirming the model is learning how to map text into a more favorable search space.

Terminology

Summary

Document expansion is recast as a document optimization problem where an instruction-tuned language model or vision language model is fine-tuned to transform documents into representations that better align with the expected query distribution under a target retriever, using GRPO with the retriever’s ranking improvements as rewards. This approach requires only black-box access to retrieval ranks and has been shown to yield retrieval gains across single-vector, multi-vector, and lexical retrievers in both code and visual document retrieval tasks.

How it works

The method formulates document optimization as a reinforcement learning problem where the policy model (the instruction-tuned LM or VLM) generates candidate transformations for a document. These transformations are evaluated counterfactually to measure their impact on retrieval quality under a fixed collection, specifically by defining the counterfactual change for a query as:

∆nDCG@k(q):= nDCG@k(q; R, D˜i) − nDCG@k(q; R, D˜)

The reward is then aggregated over positive and negative query sets associated with the document to capture retrieval-quality gains:

"ri(˜di) = 1/Q+i∑q∈Q+i∆nDCG@k(q)Q-i∑q∈Q-i∆nDCG@k(q)z>"

This process is optimized using Group Relative Policy Optimization (GRPO), which leverages relative reward comparisons between action trajectories for variance reduction without requiring an explicit value function. The policy learns to generate rewrites that are better aligned with the retriever’s representation space.

Key Components and Settings

The approach is designed to be flexible across different retrieval architectures. It applies in black-box settings and is compatible with:

  1. Single-vector retrievers (e.g., OpenAI text-embedding-3 models).

  2. Multi-vector retrievers (e.g., Jina-ColBERT-V2).

  3. Lexical retrievers (e.g., BM25).

The policy is initialized from an instruction-tuned LM or VLM, conditioned on a fixed prompt specifying the desired transformation, such as:

(For Visual Document Transformation Prompt):

Provide a comprehensive description of the document in the image in English. Begin with a summary, then follow with details. Extract all visible text and numerical values from the document.

The policy samples candidate rewrites using parameters like temperature (set to 0.7) to balance exploration and quality, while index population uses a low temperature for stable, high-quality rewrites.

Evaluation and Findings

The method was evaluated on two retrieval settings: Visual Document Retrieval (VDR) and Code Retrieval. The results consistently show that document optimization improves retrieval quality across both tasks. For example, applying document optimization to the OpenAI text-embedding-3-small model improved nDCG@5 on code from 58.7 to 66.8, slightly surpassing the more expensive OpenAI text-embedding-3-large model (66.3). Furthermore, in open-weight settings, document optimization is competitive with retriever fine-tuning, and their combination often yields the strongest performance. For instance, Joint Adaptation for Jina-ColBERT-V2 improved performance to 63.3 on VDR and from 48.6 to 61.8 on code retrieval.

Analysis of Document Transformation

An analysis of document length reveals that while zero-shot transformations substantially increase length (mean of 382 tokens), optimized documents are more compact, with a mean of 346 tokens, indicating the learned policy does not simply expand content, but selectively refines it to retain useful information. Representation space analysis shows that optimized documents form clusters and in many cases where optimization improves nDCG@5, they shift closer to the query and farther from negatives, confirming that improvements are driven by favorable relative positioning with respect to competing negatives.

Design Choices and Ablations

Key design choices were analyzed in ablations. The ranking-based reward, which uses both positive and negative queries, showed the strongest overall results because Negative queries consistently help under both ranking-based and similarity-based objectives. Regarding weak supervision for negative queries, hard negatives outperform random sampling, with the strongest results obtained using five hard negatives. The paper also considers joint retriever-policy adaptation as a form of data augmentation, where the retriever is periodically fine-tuned on the current transformed collection.

Conclusion

Document optimization provides a framework for improving retrieval by rewriting documents to better align with a target retriever, requiring only black-box access and showing gains across various architectures and tasks. The results suggest that optimizing the document space itself is a promising and underexplored direction. Future work is encouraged to extend this framework to other domains.

Improvements for AI systems

Here are specific improvements to existing retrieval and document processing AI systems based on the provided research, along with what those improved systems could achieve:


  1. A significant improvement in retrieval quality for modern neural retrievers (like bi-encoders) by optimizing the document representation itself, even when only black-box access to the retriever's ranking is available.

  2. The ability to transform source documents (text or images of documents) into retrieval surrogates—compressed, restructured, or rephrased variants—that are specifically optimized for a target retriever's representation space.

  3. Implementation of an offline Reinforcement Learning (RL) policy that fine-tunes a Language Model or Vision-Language Model (LM/VLM) to generate these retrieval-optimized document transformations using Group Relative Policy Optimization (GRPO).

  4. A mechanism for evaluating the quality of document rewrites counterfactually: by measuring the change in ranking metrics (nDCG@k) when a specific document is replaced with its rewritten version, against a fixed set of queries.

  5. The capability to improve retrieval performance on diverse tasks, specifically:

Ease retrieval performance on both Code Retrieval (rewriting code snippets into retrieval-optimized textual representations) and Visual Document Retrieval (converting PDF images into text surrogates).

  1. A system that can achieve state-of-the-art ranking gains over direct retrieval and zero-shot transformations for black-box retrievers, even when using smaller or more efficient embedding models (e.g., OpenAI text-embedding-3-small outperforms the significantly larger textembedding-3-large in some settings).

  2. The capacity to achieve competitive performance with retriever fine-tuning in white-box settings, and synergistic performance when combining document optimization with retriever fine-tuning.

  3. The ability to leverage Joint Adaptation, where the retrieval model is periodically fine-tuned on the newly transformed corpus at each iteration of the document optimization policy, leading to substantial gains (e.g., improving Jina-ColBERT-V2 from 55.8 to 63.3 on VDR).

  4. The ability to produce highly compact, high-signal document transformations that selectively refine content rather than simply expanding it, as evidenced by optimized documents being significantly smaller (mean of 346 tokens vs. zero-shot mean of 382 tokens) while maintaining or improving nDCG@5 scores.

  5. A system capable of leveraging weak supervision strategies to alleviate the high cost of annotated queries, specifically by constructing negative queries from documents that receive high similarity scores but are not labeled as relevant (hard negative mining).

These improvements enable the creation of retrieval systems that are more robust, efficient, and highly tailored to the specific needs of a target retrieval index, regardless of whether that index is a lexical search engine or a modern bi-encoder.

Sources

Related papers