Giga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text Embeddings
cs.CL
Submitted: 2026-08-24
Updated: 2026-08-24
Terminology
Sources
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation
- MMTEB: Massive Multilingual Text Embedding Benchmark
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- Distilling the Knowledge in a Neural Network
- Nomic Embed: Training a Reproducible Long Context Text Embedder
- Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation
- Mixtral of Experts
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Gecko: Versatile Text Embeddings Distilled from Large Language Models
- Gemini Embedding: Generalizable Embeddings from Gemini
- GUIDE: Guided Initialization and Distillation of Embeddings
- Representation Learning with Contrastive Predictive Coding
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Improving Text Embeddings with Large Language Models
- C-Pack: Packed Resources For General Chinese Embeddings
- One Student Knows All Experts Know: From Sparse to Dense
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering