jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers".
Jane: The paper was written by Florian Hönicke, Michael Günther, Andreas Koukounas, Mohammad Kalim Akram, Saba Sturua et al. from Jina by Elastic.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: We're starting our show with a look at a new paper called "jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers."
Jane: That title is definitely a mouthful for our listeners.
Tom: It is, but it actually describes exactly what the researchers at Jina eye are doing.
Jane: Can you explain what they mean by "geometry-preserving" in plain English?
Tom: Imagine you have a perfect map of a city where every street is in the right place.
Jane: So they want to add new landmarks without moving the existing streets?
Tom: That's a great way to put it.
Jane: And the researchers, like Florian Hönicke and Michael Günther, are trying to keep that text "map" perfect.
Lu: This approach could allow us to layer new senses onto our digital world without destroying the foundation we've already built.
Meng: I'm curious if "locked towers" means they're avoiding the hard work of actually teaching the model new things.
Lu: It's more about being smart with what we've already learned, Meng.
Lalam: It creates a unified way for machines to experience the rhythm of speech and the color of images simultaneously.
Jane: The "locked" part refers to the fact that they keep the vision and audio encoders frozen.
Tom: They aren't changing the internal workings of those specialized encoders at all.
Jane: They only train the connection between those encoders and the text model.
Tom: That's where the "alignment" part of the title comes in.
Lu: It's a brilliant way to respect the specialized knowledge already inside those pre-trained models.
Meng: I see the appeal, but does keeping them "locked" limit how well they can learn to talk to the text model?
Lu: It actually prevents the new information from corrupting the existing text knowledge.
Jane: It's like building a new bridge to an existing city instead of tearing the city down to build a new road.
Tom: That stability is why they focused so much on preserving the geometry.
Lalam: This allows the digital world to expand its senses without losing its original voice.
Jane: It sounds like they are prioritizing compatibility over everything else.
Tom: Let's see how they actually pull off that feat in the summary.
Summary: Tom: We've touched on the concept, but now we need to talk about how the GELATO method actually works.
Jane: The paper explains that they use these small "projectors" to bridge the gap between modalities.
Tom: They aren't just using any encoders, either.
Jane: They're using Qwen3 point 5 for vision and Qwen2 point 5-Omni for audio.
Tom: These are already very smart, language-aligned encoders.
Jane: So the projector just has to translate their output into the text model's hidden space?
Tom: Precisely, and they've made it incredibly efficient.
Meng: How much of the model are we actually talking about when we say "efficient"?
Tom: They only trained about zero point three five percent of the total weights.
Meng: That's an incredibly small fraction of the parameters.
Lu: It's a very surgical approach to machine learning.
Jane: They also use these special "delimiter tokens" to tell the model where an image or an audio clip starts and ends.
Tom: It's like putting brackets around a quote in a sentence.
Lu: It allows the model to process a sequence that might have text, then an image, then more text.
Meng: Does this mean they can handle video too?
Tom: They do, by treating a video as a sequence of sampled frames.
Jane: It's a very modular way of thinking about data.
Lalam: It's how we move toward a truly omni-modal understanding of human culture.
Tom: But does this efficient approach actually deliver high-quality results?
Improvements: Tom: The results in "jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers" are quite striking.
Jane: I was particularly struck by the ViDoRe scores for document retrieval.
Tom: The nano version actually hit seventy-nine point two five.
Jane: And it did that with only zero point three one billion active parameters.
Meng: That's much more efficient than the LCO-Omni-7B model, which uses nearly nine billion parameters.
Lu: But we have to be honest about the video performance.
Tom: You're right, the video scores are definitely the weak point here.
Jane: They also looked at the "modality gap" using UMAP plots.
Tom: Those plots show how the different types of data live in the same space.
Jane: In some models, text and images live in completely separate clusters.
Tom: But in the Jina models, they're all interleaved and mixed together.
Lu: That's the "geometry-preserving" part in action.
Meng: I noticed they also did an ablation on whether to unfreeze the encoders.
Tom: Yeah, they found that unfreezing the vision encoder actually made things worse.
Jane: It destabilized the whole system because the projector wasn't ready yet.
Meng: That makes sense; you don't want to change the foundation while you're still building the walls.
Lu: They did find that a two-stage process for audio might be a way to improve things later.
Jane: And we can't forget the Matryoshka feature, which lets you shrink the embeddings.
Tom: It keeps the important information in the first few dimensions.
Jane: It's a very robust way to handle data compression.
Tom: It seems like they've found a very stable way to scale up.
Conclusion: Tom: We're wrapping up our look at "jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers."
Jane: It's been a fascinating look at how to expand eye's senses efficiently.
Tom: They've shown that you don't need to retrain the whole world to add new modalities.
Lu: This modularity is going to change how we approach multimodal research.
Meng: It definitely makes the engineering side of things a lot more manageable.
Lalam: It brings us closer to a digital intelligence that perceives the world as a single, cohesive experience.
Tom: Thanks for joining us, we'll see you at the next paper.
Florian Hönicke, Michael Günther, Andreas Koukounas, Mohammad Kalim Akram, Saba Sturua, Han Xiao
Jina by Elastic
cs.CL
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 11 pages, 9 figures, 5 tables
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 82/100
The gist: The paper introduces GELATO (Geometry-preserving Embeddings via Locked Aligned TOwers), a novel approach to constructing multimodal embedding models.
Key concepts
- Geometry-preserving
- This refers to keeping the existing "map" of text knowledge intact while adding new senses like vision or audio. It ensures that adding new information does not corrupt or move the established relationships within the original text model, maintaining the foundation's stability.
- Locked Aligned Towers
- A method where specialized vision and audio encoders are kept frozen during training. Instead of retraining the whole model, researchers only train the connection, or "projector," between these encoders and the text model to align them without destroying existing knowledge.
- Matryoshka feature
- A technique that allows embeddings to be shrunk while keeping the most important information in the first few dimensions. This provides a robust way to handle data compression, making it easier to manage and process large amounts of data.
Terminology
Summary
The paper introduces GELATO (Geometry-preserving Embeddings via Locked Aligned TOwers), a novel approach to constructing multimodal embedding models. The authors state: We build on the VLM-style architecture, in which non-text encoders are adapted to produce input for a language model, which in turn generates embeddings for all varieties of input.
The result is the jina-embeddings-v5-omni suite, a pair of models that encode text, image, audio, and video input into a single semantic embedding space.
The two models are:
-
jina-embeddings-v5-omni-nano, based on jina-embeddings-v5-text-nano with 0.24B parameters in its base text-only model
-
jina-embeddings-v5-omni-small, based on jina-embeddings-v5-text-small with 0.67B parameters
The core idea is that "GELATO extends the two Jina Embeddings v5 Text models to support additional modality by adding encoders for images and audio. The backbone text embedding models and the added non-text modality encoders remain frozen. We only trained the connecting components, representing 0.35% of the total weights of the joint model."
The architecture integrates:
-
Vision encoders from Qwen3.5-2B and Qwen3.5-0.8B, which have been adapted from SigLIP2 So400m and SigLIP2 Base respectively
-
The Qwen2.5-Omni audio encoder, which has been adapted from Whisper-large-v3
The paper lists four main contributions:
-
We describe GELATO and apply it in the construction of the jina-embeddings-v5-omni model suite by extending the Jina Embeddings v5 Text suite to support other media.
-
"We contribute to the open embedding ecosystem by releasing the jina-embeddings-v5-omni model collection, comprising two base models and eight task-specific variants for retrieval, classification, clustering, and text-matching across Small and Nano scales."
-
We evaluate jina-embeddings-v5-omni and comparable models across a range of standard benchmarks, and show that GELATO produces competitive results.
-
We analyze the design rules behind GELATO through ablations on projector training, encoder choice, and Matryoshka truncation, and separately quantify training efficiency.
Regarding input sequence construction: Each input is serialized as one sequence of tokens. Text remains ordinary text tokens; non-text modalities are represented by placeholder runs inside modality delimiters.
An image is encoded as × N with N visual slots. An audio input is encoded as × K with K audio slots. A video is a concatenation of one visual segment per sampled frame, with audio preceding frames if present.
Regarding projectors: "jina-embeddings-v5-omni uses image and audio encoders extracted from Qwen3.5 and Qwen2.5-Omni, respectively. Because their output dimensions do not match Jina Embeddings v5 Text's input, we replace the source projection layers with new projectors that map into the text hidden space. For vision, the projector applies
LayerNorm, a 2×2 spatial merge, fc vision 1, GELU, and fc vision 2. Only fc vision 2 is trained, mapping
4096→1024 for Small and
3072→768 for Nano. For audio, a
randomly-initialized fc audio layer projects
the encoder's native 1280 dimension output into jina-embeddings-v5-omni-small's 1024-dimension input space and jina-embeddings-v5-omni-nano's 768-dimension one."
The trainable set consists of fc vision 2, fc audio, and the modality-delimiter embeddings.
The paper notes: jina-embeddings-v5-omni-small learns the vision and audio start/end delimiter embeddings used in Section 3.2; jina-embeddings-v5-omni-nano learns only the audio start/end delimiter embeddings.
Training uses bidirectional in-batch InfoNCE with Matryoshka representation learning
with temperature tau = 0.02, summing loss over Matryoshka prefix dimensions KSmall = 32, 64, 128, 256, 512, 768, 1024 and KNano = 32, 64, 128, 256, 512, 768. The optimizer is AdamW with beta1=0.9, beta2=0.999, weight decay 0.01, learning rate 2·10−4 with 500 linear warmup steps, bf16 mixed precision, distributed data parallelism across 4 NVIDIA H100 GPUs, global batch size 256, and 15,000 optimizer steps per run. This yields 2 × 4 × 2 = 16 projector-training runs in total.
The training mixture is described: The mixture is full of text-rich and complex images like scans and diagrams, matching practical enterprise search and RAG systems that operate over real-world multimodal documents.
Key results from Table 1 show:
-
jina-embeddings-v5-omni-small achieves a 54.04 four-modality average,
slightly above LCO-Embedding-Omni-3B (53.83) and below only the larger LCO-Embedding-Omni-7B score of 54.43
-
jina-embeddings-v5-omni-small has
the strongest text-only performance and the best overall score among models below 5B parameters
-
Video performance
lags significantly compared to the baseline models
Table 2 shows strong visual document retrieval: "jina-embeddings-v5-omni-small scores 79.25 with 0.92B active text+image-path parameters, above LCO-Embedding-Omni-3B (78.24) and close to LCO-Embedding-Omni-7B (80.32). jina-embeddings-v5-omni-nano matches that with 79.25 at 0.31B active parameters."
On modality geometry (Section 5.2), the paper reports on MS-COCO Karpathy split and Clotho v2: Our jina-embeddings-v5-omni-small variant (1.57B, frozen-tower) sits in third place at 68.0% / 57.0%
for image-text retrieval, and reaches 16.3% / 15.2%
for audio-text retrieval. The paper notes: the gap between the jina-embeddings-v5-omni-small variant and LCO-Omni-7B is markedly larger on the audio pair (∼11–15% absolute) than on the image pair (∼6–7%).
Figure 4 shows UMAP geometry: "LanguageBind's frozen per-modality towers produce the canonical modality-gap pattern; the unified-decoder omni models, including the frozen-tower jina-embeddings-v5-omni-small and jina-embeddings-v5-omni-nano, produce interleaved geometry."
Ablation studies (Section 6) test which layers to train:
-
Vision:
The fc vision 2-only recipe (I) is sufficient: it reaches 0.158, while training fc vision 1 from the start (II) ends slightly lower at 0.153. Unfreezing the encoder from step 0 (III) is clearly harmful, ending at 0.079.
-
Audio:
The fc audio-only recipe (I) is sufficient for this budget: it reaches 0.398, while unfreezing the audio encoder from step 0 (II) ends lower at 0.367.
A two-stage continuation (III) reaches 0.419,an absolute gain of 0.022 over I.
Matryoshka truncation tests (Figure 9) show: Image embeddings behave similarly to text ones... Audio also preserves most of its score at 256 dimensions, while video degrades much more heavily at small dimensions.
The text and image curves overlap almost exactly at every truncation level for a given model size.
Training efficiency (Table 5) shows: projector training makes vision runs 1.8× faster and audio runs 3.2–3.9× faster at the 15k-step budget, with lower peak GPU memory in every case.
The paper concludes: jina-embeddings-v5-omni-small is the best-performing open-weight embedding model below 2B parameters that supports text, audio, images, and video.
The authors note that the ablations suggest that projector-only alignment can serve as a compatibility-preserving initialization for rich multimodal training,
and identify future work including the choice of non-text encoders
and jointly training projectors for multiple modalities together.
They also acknowledge: overall video performance remains weak. We hope to improve performance in this area in future models.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Add a frozen-tower projector layer (fc vision 2 and fc audio) to any existing text embedding model, enabling it to process images, audio, and video without retraining the backbone.
What the improved system can do:
-
Accept image inputs (screenshots, documents, infographics) and encode them into the same semantic space as text
-
Accept audio inputs (speech, music, environmental sounds) and encode them for retrieval
-
Accept video inputs (sampled frames + audio track) as a single embedding
-
Maintain bit-identical text performance while adding new modalities
Improvement: Implement dynamic weight loading that selects task-specific LoRA adapters and modality projectors based on the downstream task (retrieval, classification, clustering, text-matching).
Improvement: Apply the Matryoshka truncation schedule (dimensions 32, 64, 128, 256, 512, 768, 1024) to multimodal embeddings during training.
Improvement: Use the GELATO training recipe: freeze all encoders and text backbone, train only the second vision projector layer (fc vision 2) and audio projector (fc audio) with bidirectional InfoNCE loss.
Improvement: Apply the geometry analysis from Section 5.2 to detect and mitigate modality gaps in embedding spaces.
Improvement: Implement the two-stage training for audio (Stage I: train fc audio only; Stage II: unfreeze audio encoder with 20× reduced learning rate).
Improvement: Leverage the frozen text backbone's multilingual capabilities through the projector architecture.
Improvement: Use the vision encoder's ability to process text-rich images (scans, diagrams, charts) through the frozen Qwen3.5 vision tower.
Improvement: Implement the input sequence construction with configurable visual/audio slots (N visual slots, K audio slots).
Improvement: Use the performance profile from Table 4 to guide model selection for specific use cases.
Abstract
In this work, we introduce GELATO (Geometry-preserving Embeddings via Locked Aligned TOwers), a novel approach to multimodal embedding models. We build on the VLM-style architecture, in which non-text encoders are adapted to produce input for a language model, which in turn generates embeddings for all varieties of input. We present the result: the jina-embeddings-v5-omni suite, a pair of models that encode text, image, audio, and video input into a single semantic embedding space. GELATO extends the two Jina Embeddings v5 Text models to support additional modality by adding encoders for images and audio. The backbone text embedding models and the added non-text modality encoders remain frozen. We only trained the connecting components, representing 0.35% of the total weights of the joint model. Training is therefore much more efficient than full-parameter retraining. Additionally, the language model remains effectively unaltered, producing exactly the same embeddings for text inputs as the Jina Embeddings v5 Text models. Our evaluations show that GELATO produces results that are competitive with the state-of-the-art, yielding nearly equal performance to larger multimodal embedding models.
Sources
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation
- Qwen3-VL Technical Report
- e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
- Enhancing Vision Foundation Models via Multimodal Continual Pre-Training
- Qwen2.5-Omni Technical Report
- MAEB: Massive Audio Embedding Benchmark
- MMTEB: Massive Multilingual Text Embedding Benchmark
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- E5-V: Universal Embeddings with Multimodal Large Language Models
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
- jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
- Jina CLIP: Your CLIP Model Is Also Your Text Retriever
- A resource- and computationally-efficient protocol for multipartite entanglement distribution in Bell-pair networks
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
- ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval
- Nomic Embed Vision: Expanding the Latent Space
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Multilingual E5 Text Embeddings: A Technical Report
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering