UEmbed: Unified Sparse and Dense Multimodal Embeddings
Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Zhijie Nie, Yilun Zhao, Shu Wu
cs.CV, cs.AI, cs.CL, cs.IR
Submitted: 2026-08-03
Code: https://github.com/Alibaba-NLP/DeepResearch
Project page: https://alibaba-nlp.github.io/UEmbed
License: http://creativecommons.org/licenses/by/4.0/
The gist: Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation.
Terminology
Abstract
Sparse retrieval underpins modern search systems, from web search to retrieval-augmented generation. Existing work has introduced Learned Sparse Retrieval (LSR) to push beyond exact lexical matching toward richer semantics. Yet LSR has so far remained tied to encoder-style bidirectional architectures, and its extension to multimodal settings still relies heavily on auxiliary cross-modal modules. To address these limitations, we introduce UEmbed (Unified Embedding), a decoder-only multimodal embedding model that produces both sparse lexical and dense representations in one causal forward pass. UEmbed appends N learnable special tokens to the input and partitions the vocabulary into N disjoint subsets. Each token's causal hidden state predicts sparse weights over its assigned subset, and the N subsets are concatenated into the full sparse vector. Trained on public data, we release UEmbed at 2B, 4B, and 9B scales. UEmbed-9B reaches 71.8 (dense) and 71.0 (sparse) on MMEB-v2, outperforming multimodal embedding models trained on publicly available data (e.g., RzenEmbed). On BEIR, UEmbed also remains competitive with strong dense and sparse baselines. Furthermore, we demonstrate the practical utility of UEmbed across three dimensions: effectiveness, efficiency, and agentic applications. Overall, UEmbed offers a new paradigm: it unifies dense and sparse embeddings in one model, while further extending sparse retrieval to unify text and multimodal inputs.
Sources
- Qwen3-VL Technical Report
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- Think Then Embed: Generative Context Improves Multimodal Embedding
- Mistral-SPLADE: LLMs for better Learned Sparse Retrieval
- SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval
- Learning Retrieval Models with Sparse Autoencoders
- Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
- LoRA: Low-Rank Adaptation of Large Language Models
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
- SPLADE-v3: New baselines for SPLADE
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
- Revisiting Text Ranking in Deep Research
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- Minimizing FLOPs to Learn Efficient Sparse Representations
- OpenAI GPT-5 System Card
- LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum
- GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models