GEM: A Generative Embedding Model Bridging Reasoning and Retrieval

arXiv:2608.13200 · cs.CL, cs.AI, cs.IR · Submitted 2026-08-14 · Read on arXiv

Zhili Shen, Craig Macdonald

University of Glasgow

cs.CL, cs.AI, cs.IR

Submitted: 2026-08-14

Updated: 2026-08-17

Code: https://github.com/huggingface/acceleratehttps:

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: Based on the paper, here is a detailed summary: GEM: A Generative Embedding Model Bridging Reasoning and Retrieval This paper introduces GEM, a generative embedding model that unifies reasoning and

Terminology

Summary

Based on the paper, here is a detailed summary:

GEM: A Generative Embedding Model Bridging Reasoning and Retrieval

This paper introduces GEM, a generative embedding model that unifies reasoning and retrieval within a single model. The core problem addressed is the growing gap between how users express complex information needs and how conventional retrievers interpret them, which often rely on surface-level lexical or semantic matching.

Core Approach: Generate-Then-Encode Paradigm

GEM integrates reasoning and embedding by first reasoning over a query and then encoding the enriched context for retrieval. Specifically, given a query, GEM constructs a prompt that instructs it to reason about user intent and relevance criteria. It then generates a response, and for retrieval, it learns the similarity between a document and the concatenated prompt-response pair. An embedding token is appended to the end of the response to produce the embedding, and its representation is computed efficiently by reusing the KV cache from generation.

Key Contributions and Methodology

  1. Unified Model: GEM is jointly trained with contrastive and causal language modelling objectives, preserving generation while enabling effective embedding. This addresses the issue of catastrophic forgetting in LLM-based embedding models. The final training objective is a weighted sum of generation and embedding losses: LGEM = λgen Lgen + λemb Lemb.

  2. Tailored Data Generation: To align embedding with generation, the authors introduce a data synthesis strategy that constructs document pairs conditioned on validated reasoning. This involves:

  • Response Generation and Filtering: Sampling candidate responses from an LLM and filtering them using an LLM-based relevance classifier to ensure the original positive document remains relevant under the generated reasoning.

  • Document Generation: Generating positive and hard negative documents conditioned on the reasoning. Hard negatives share similar topics but contain subtle contradictions, discouraging simple surface-level matching.

  1. Training: GEM is trained from Qwen3-4B-Instruct-2507 for 500 steps with an effective batch size of 512. The training data is augmented from Promptriever and ReasonIR datasets, consisting of 370K samples.

Experimental Results and Research Questions

The paper evaluates GEM on reasoning-intensive (BRIGHT) and instruction-following (FollowIR, InstructIR) retrieval tasks, addressing several research questions:

  • RQ1 (Reasoning-intensive retrieval): GEM achieves an average nDCG@10 of 29.1 on BRIGHT, outperforming single-model baselines like GritLM-7B and ReasonIR-8B. It excels at theorem-based tasks, improving the average nDCG@10 over its embedding-only variant from 19.8 to 32.0.

  • RQ2 (Test-time compute scaling): GEM's retrieval performance can be improved by prompting it to generate longer responses. The average nDCG@10 on BRIGHT peaks at 30.1 when prompted with n=1024 words. Encoding time remains stable when reusing the KV cache, unlike pipelines with independent models.

  • RQ3 (Instruction-following retrieval): GEM performs on par with Promptriever and outperforms other LLM-based retrievers on FollowIR, achieving a p-MRR of +11.7. It shows consistent improvements over its embedding-only variant, with notable gains in p-MRR (+6.8 → +11.7) and Robustness@10 (46.2 → 54.8).

  • RQ4 (Comparison with query expansion): LLM-based query expansion methods like HyDE and Query2Doc are insufficient to achieve strong instruction-following retrieval performance like GEM. They often degrade p-MRR for models like Promptriever.

  • RQ5 (Trade-off of unifying generation and embedding): The unified model shows a slight degradation in embedding performance compared to an embedding-only variant trained on the same data (p-MRR +12.5 → +11.7 on FollowIR, nDCG@10 30.0 → 29.1 on BRIGHT), but it remains strong on generative tasks.

  • RQ6 (Effects of training data components): Incorporating hard queries from ReasonIR substantially improves p-MRR on FollowIR (+8.5 → +11.7). Adding non-reasoning original samples regularises GEM, mitigating overfitting. Disabling document generation substantially reduces nDCG@10 on BRIGHT (29.1 → 25.8).

Conclusion and Limitations

GEM demonstrates strong performance on reasoning-intensive and instruction-following retrieval tasks, despite being a 4B-parameter model trained with substantially less compute than larger baselines. Its generative nature allows test-time compute scaling through prompting.

The paper acknowledges limitations, including:

  • Inability to replicate experiments with larger backbones due to limited computing resources.

  • Potential hallucinations during document generation and inference.

  • The cost of autoregressive generation, although GEM saves query-side encoding time by reusing the KV cache.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:

1. Unified Reasoning-Retrieval Architecture

  • Improvement: Integrate a generate-then-encode paradigm into a single model, where the model first reasons about user intent and relevance criteria, then encodes the enriched context for retrieval. Use a shared KV cache to avoid redundant computation.

  • Improved capability: The AI system can handle complex, multi-hop queries (e.g., Find papers that contradict the methodology of Smith et al. 2020) by explicitly reasoning about what constitutes a relevant document, rather than relying on surface-level keyword or semantic similarity. It can do this in one forward pass, with no separate retrieval pipeline.

2. Test-Time Compute Scaling for Retrieval

  • Improvement: Allow the model to generate longer reasoning responses at inference time (e.g., prompt it to produce 1024-word analyses) without increasing encoding time, by reusing the KV cache from generation.

  • Improved capability: The system can dynamically trade off latency for accuracy. For high-stakes queries (e.g., legal or medical retrieval), it can generate more detailed reasoning to improve nDCG@10 by up to 10% (from 29.1 to 30.1 on BRIGHT), while keeping query-side encoding time constant. For low-latency applications, it can use shorter responses.

3. Reasoning-Conditioned Data Synthesis

  • Improvement: Implement a data generation pipeline that (a) samples candidate reasoning responses, (b) filters them using an LLM-based relevance classifier to ensure the positive document remains relevant, and (c) generates hard negative documents with subtle contradictions to the reasoning.

  • Improved capability: The AI system can be trained on high-quality, reasoning-aligned data, making it robust to adversarial or nuanced queries. It will avoid overfitting to surface-level patterns and instead learn to match documents based on deeper logical consistency, as evidenced by a 12.8% improvement on theorem-based tasks (nDCG@10 from 19.8 to 32.0).

4. Joint Training with Contrastive and Causal Objectives

  • Improvement: Train the model with a weighted sum of generation loss (Lgen) and embedding loss (Lemb), preserving the model's generative abilities while optimizing for retrieval. Use a ratio (λgen/λemb) that balances both.

  • Improved capability: The system can simultaneously perform generative tasks (e.g., summarization, question answering) and retrieval tasks without catastrophic forgetting. This enables a single model to serve as both a retriever and a generator, reducing infrastructure complexity and enabling end-to-end pipelines (e.g., retrieve relevant documents, then generate an answer from them).

5. Instruction-Following Retrieval with Hard Query Augmentation

  • Improvement: Incorporate hard, reasoning-intensive queries from datasets like ReasonIR into training data, alongside original non-reasoning samples to regularize the model.

  • Improved capability: The system can accurately follow complex user instructions in retrieval, such as Find documents that support the claim X but not Y or Retrieve sources that are from peer-reviewed journals only. This improves p-MRR by 37% (from +8.5 to +11.7 on FollowIR) and robustness (Robustness@10 from 46.2 to 54.8), making it reliable for real-world information-seeking tasks where users express nuanced needs.

6. KV-Cache Reuse for Efficient Multi-Turn Retrieval

  • Improvement: During inference, reuse the KV cache from the generation phase to compute the embedding token's representation, avoiding redundant forward passes.

  • Improved capability: The system can handle iterative retrieval scenarios (e.g., conversational search) where a user refines their query based on initial results. It can re-encode the updated query with minimal additional compute, enabling real-time interactive retrieval without latency spikes.

7. Mitigation of Hallucination in Reasoning

  • Improvement: Use the LLM-based relevance filter during training to discard reasoning responses that would misalign with the positive document, and generate hard negatives that explicitly contradict the reasoning to teach the model to avoid false positives.

  • Improved capability: The system will produce more trustworthy retrieval results, especially for queries requiring logical inference. It will be less likely to retrieve documents that merely mention the same keywords but contradict the user's implied criteria, reducing false positives in high-precision domains like fact-checking or academic literature review.

Sources

Related papers