Douyin Multimodal Embedding Model Technical Report
Haonan Chen, Chu Li, Zhicheng Wang, Yuanwei Liu, Yuanjiang Wang, Shaohua Jiang, Zhicheng Dou
cs.IR, cs.CL, cs.CV
Submitted: 2026-08-18
Updated: 2026-08-19
Comments: Technical Report
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal representation learning is a cornerstone of modern AI.
Terminology
Abstract
Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with complex modalities and massive-scale content, such as Douyin, Xiaohongshu, and YouTube, demand both efficiency under billion-scale indexing and fine-grained discrimination for hard matching. Existing MLLM embedding models rarely satisfy both. Contrastive models are efficient but rely on pair-level supervision too coarse for fine-grained distinctions, while CoT-based models improve discrimination through explicit generation impractical to serve online. We present Douyin Multimodal Embedding (DME), a model trained in two stages to combine both strengths. Stage 1 performs large-scale contrastive pre-training that establishes a unified multimodal embedding space with broad modality and task coverage. Stage 2 supplements semantic sufficiency, the property that an embedding is grounded in retrieval-relevant evidence and preserves fine-grained counterpart-side semantics, via two mechanisms. Evidence-Grounded Typed Latent Reasoning organizes retrieval evidence through hidden-space latent reasoning, and Cross-Conditional Reconstruction enforces counterpart-side semantics through cross-directional autoregressive reconstruction. Both act only during training and add only marginal query-side overhead, so DME serves as efficiently as a standard contrastive encoder. On MMEB-v2, DME reaches state-of-the-art results at comparable scales for its 2B and 9B variants (74.8 and 78.4), with especially strong video and visual-document tasks. In production, DME delivers a 2.92% relative gain on Douyin's in-house offline evaluation set, is deployed across Douyin scenarios such as generative, image, and AI search, and yields a 0.1% Lifetime (LT) gain in online A/B testing on Douyin search.
Sources
- MM-BRIGHT: A Multi-Task Multimodal Benchmark for Reasoning-Intensive Retrieval
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
- Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning
- Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search
- Think Then Embed: Generative Context Improves Multimodal Embedding
- Reason to Contrast: A Cascaded Multimodal Retrieval Framework
- DeepSeek-V3 Technical Report
- MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising
- TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
- E5-V: Universal Embeddings with Multimodal Large Language Models
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
- U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG