Generative Universal Multimodal Retrieval with Dual-role Identifiers

arXiv:2608.12987 · cs.IR, cs.AI · Submitted 2026-08-13 · Read on arXiv

Kaipeng Li, Haitao Yu, Xuanchen Zhou

Independent Researcher · University of Tsukuba · University of Tsukuba

cs.IR, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: This paper is under review

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper proposes DrIG, a novel generative framework for universal multimodal retrieval featuring dual-role identifiers.

Terminology

Summary

The paper proposes DrIG, a novel generative framework for universal multimodal retrieval featuring dual-role identifiers. The work addresses three key challenges in generative information retrieval (GIR): (1) constrained left-to-right decoding is vulnerable to prefix-level errors and local optima; (2) most prior GIR research remains largely unimodal, leaving instruction-aware retrieval across text, image, and mixed image-text items underexplored; and (3) discrete identifier-based GIR offers higher efficiency but its retrieval accuracy lags behind cutting-edge dense-vector-based retrieval methods.

The central innovation is assigning each candidate a single residual-quantized identifier that serves two complementary roles:

  • Sequential role: The identifier is decoded autoregressively, where the first token explicitly models modality (image, text, or image-text pair) and remaining tokens capture progressively finer semantics.

  • Set-based role: The same tokens are reinterpreted as an unordered set to provide a prefix-independent relevance prior, which guides constrained beam search and alleviates local-optimum errors.

The framework employs a large multimodal model (LMM) - specifically Qwen2-VL - to encode queries and candidates into a shared embedding space. Following the explicit one-word limitation (EOL) strategy, the model is prompted to summarize inputs in one word. A two-stage fine-tuning strategy is deployed: first adapting the LMM for text-to-text retrieval using natural language inference data, then fine-tuning on multimodal retrieval tasks using contrastive learning with InfoNCE loss.

  • Sequential role: Uses residual quantization (RQ) with a first-level modality codebook of size 3 (image, text, image-text pairs), followed by semantic residual codebooks. The training objective combines residual quantization loss, contrastive loss, and MSE loss.

  • Set-based role: Reinterprets the sequential identifier as an unordered set. A token-level score vector is computed via a query-to-token mapping, and order-invariant relevance scores are aggregated. Training uses contrastive and margin ranking losses.

  • Inference: Combines sequential decoding scores with set-based global relevance priors during Trie-constrained beam search. The unified expansion score is: f(t≤i; zq) = δ(t≤i) + η(t<i; zq) + E i dec[t i]·h i + λφ(t≤i; zq)

  • Query augmentation: Uses query-target interpolation in continuous embedding space with Beta-distributed interpolation coefficients.

  • Training objectives: Combines generative cross-entropy loss with a discriminative ranking objective using adaptive margins derived from teacher similarity signals.

Combines generative retrieval with dense vector-based reranking: the generative retriever produces a compact top-k candidate list, which is then reranked using cosine similarity in the original embedding space. This preserves scalability while recovering fine-grained distinctions lost during quantization.

  • Local-pool retrieval: DrIG improves average score from 29.5 (GENIUS) to 38.0, a relative gain of 28.8%. With dense reranking, DrIG-C and DrIG-LT achieve 48.7 and 50.4 respectively.

  • Global-pool retrieval: DrIG improves from 28.6 to 36.4 (27.3% relative gain). DrIG-C and DrIG-LT reach 47.1 and 48.9.

  • DrIG consistently outperforms GENIUS across all retrieval tasks, with particularly large gains on knowledge-intensive tasks (e.g., InfoSeek in Task 6: 11.4→25.0 in local-pool setting).

On Flickr30K and MSCOCO:

  • M-BEIR-trained DrIG achieves strong zero-shot performance (Flickr30K: 59.0/83.1/88.2 for R@1/R@5/R@10)

  • DrIG-LT with LamRA reranking reaches 75.8/90.0/91.6 on Flickr30K, surpassing the strongest prior hybrid baseline ComGTIR-DHclip

  • In-domain training provides additional gains (Flickr30K: 76.9/92.5/94.8 for DrIG-LT)

Key findings from Table 5:

  • Contrastive loss before quantization is essential: Removing Lsr cl causes performance collapse (MSCOCO R@1 drops from 41.8 to 0.5)

  • Trie constraint is critical: Removing it causes large drops (MSCOCO R@10: 79.4→21.6)

  • Set-based role provides consistent but moderate gains across tasks

  • Query augmentation contributes substantially to decoder robustness

  • Discriminative ranking loss improves ranking consistency

  • Modality codebook is beneficial for heterogeneous retrieval tasks

  • Generative retrieval methods (GENIUS, DrIG) maintain nearly flat QPS as candidate pools grow from 5K to 300K

  • Dense retrieval baselines show decreasing throughput with larger pools

  • DrIG-LT substantially improves effectiveness (36.4→48.9 average global-pool score) while preserving scalability advantages

  • Codebook configuration: Increasing quantization depth (L from 4 to 8) brings consistent improvement; 8×4096 is the default

  • Beam size: Performance improves rapidly from beam 1 to 20, then saturates; QPS decreases from 972.6 (beam 1) to 37.6 (beam 50)

  • Reranking depth: Most beneficial when k increases from very small to moderate range; k=50 used for maximum effectiveness

  • Decoder backbone: T5-small (30M params) provides the most balanced cross-task performance; larger decoders help text-centric tasks but hurt visually grounded tasks

  • Global prior weight λ: λ=1.0 provides strong and stable performance across most settings

The paper demonstrates that dual-role identifiers provide a promising direction for scalable generative multimodal retrieval. The hybrid approach combining generative retrieval with dense reranking narrows the effectiveness gap between generative and dense paradigms while preserving scalability. The authors identify several future directions: developing more expressive identifiers, exploring end-to-end optimization, supporting dynamic candidate collections, and evaluating on larger-scale scenarios including video retrieval and retrieval-augmented multimodal generation.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

  • Improvement: Integrate the dual-role identifier framework into my retrieval pipeline, using a first token that explicitly encodes modality (text, image, or mixed) followed by hierarchical semantic tokens via residual quantization.

  • Capability: I can now retrieve across heterogeneous item types (pure text, pure image, image-text pairs) in a single generative pass, without requiring separate retrieval models per modality. This enables unified search over mixed digital libraries, e-commerce catalogs, or social media feeds.

  • Improvement: Adopt the set-based reinterpretation of identifiers—treating the token sequence as an unordered set to compute a global relevance prior that is independent of partial decoding order. This guides beam search even when early tokens are suboptimal.

  • Capability: I can avoid cascading errors from early mispredictions in autoregressive generation. For example, when generating a product recommendation, if the first semantic token is slightly off, the set-based prior still steers the final output toward the correct item, improving accuracy on long-tail or knowledge-intensive queries.

  • Improvement: Combine my generative retrieval (fast, scalable, discrete) with dense vector reranking on the original continuous embeddings, using the generative model to produce a compact top-k candidate set before fine-grained cosine similarity reranking.

  • Capability: I can handle candidate pools up to 300K items with near-constant query-per-second throughput, while recovering fine-grained distinctions lost during quantization. This is ideal for real-time search in large-scale document stores, where I first generate a shortlist and then rerank for precision.

  • Improvement: Implement the two-stage training strategy: (1) adapt the base LMM for text-to-text retrieval using natural language inference data, then (2) fine-tune on multimodal retrieval with InfoNCE contrastive loss before quantization is applied.

  • Capability: I can leverage existing text-retrieval knowledge to bootstrap multimodal understanding, and the contrastive pre-training ensures that the embedding space is well-structured before discrete identifiers are assigned. This prevents catastrophic performance collapse (as seen when contrastive loss is removed) and improves zero-shot transfer to unseen retrieval tasks.

  • Improvement: Apply Beta-distributed interpolation between query and target embeddings during training to augment query diversity and improve decoder robustness.

  • Capability: I can generalize better to paraphrased or noisy queries. For instance, if a user asks photos of red cars vs. pictures of crimson automobiles, the augmented training helps my generative decoder map both to the same identifier, improving recall on varied user phrasings.

  • Improvement: Use teacher-similarity signals to set adaptive margins in the ranking loss, rather than fixed margins, during joint generative-discriminative training.

  • Capability: I can produce more consistent rankings where the margin between relevant and irrelevant items scales with their true semantic distance. This improves top-1 accuracy in tasks like image caption retrieval, where some negatives are harder to distinguish than others.

  • Improvement: Use a first-level codebook of size 3 (image, text, mixed) followed by semantic residual codebooks, allowing different quantization depths or codebook sizes per modality if needed.

  • Capability: I can allocate more identifier capacity to visually complex items (e.g., scene graphs) while using compact codes for simple text, optimizing both accuracy and storage efficiency in multimodal databases.

  • Improvement: Adopt the generative retrieval paradigm where identifiers are decoded without scanning the entire candidate pool, enabling near-constant inference time regardless of collection size.

  • Capability: I can support dynamic, growing collections (e.g., continuously updated news archives or user-generated content) without re-indexing or retraining, since the generative decoder maps queries to identifiers directly. This is a key advantage over dense retrieval, which requires recomputing embeddings for new items.

Sources

Related papers