TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens
cs.AI
Submitted: 2026-05-15
Updated: 2026-09-19
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning.
Terminology
Abstract
Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this paradigm, a generative model produces explicit reasoning traces for a multimodal query, with the final representation extracted from an <eos> embedding token attending to both the query and the reasoning. Despite its effectiveness, the computational overhead of generating explicit CoT traces is often prohibitive. In this work, we propose replacing explicit CoT with latent think tokens, which are interpreted as latent variables that can produce explicit CoT traces as observed variables. By optimizing think tokens using CoT generation loss and subsequent embedding tokens using contrastive loss, we produce high-performance, reasoning-aware representations at a constant inference cost. Our study investigates two key architectural designs: 1) how think and embeddings tokens should be extracted from the same LLM backbone. 2) how the tokens should be trained as two dependent tasks. We introduce TTE-Flash-2B, a reasoning-aware multimodal representation model that outperforms its explicit-CoT counterpart on the MMEB-v2 benchmark, while producing latent think tokens that are interpretable both textually and visually. Furthermore, zero-shot evaluation across 15 video datasets reveals scaling behavior as the number of think tokens increases.
Sources
- Think Then Embed: Generative Context Improves Multimodal Embedding
- Vision Transformers Need Registers
- Implicit Chain of Thought Reasoning via Knowledge Distillation
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- ColPali: Efficient Document Retrieval with Vision Language Models
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Think before you speak: Training Language Models With Pause Tokens
- Training Large Language Models to Reason in a Continuous Latent Space
- PLUME: Latent Reasoning Based Universal Multimodal Embedding
- Categorical Reparameterization with Gumbel-Softmax
- E5-V: Universal Embeddings with Multimodal Large Language Models
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- Auto-Encoding Variational Bayes
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
- MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control
- Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection