SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia
cs.CL
Submitted: 2026-06-02
Updated: 2026-09-09
Comments: Accepted to EMNLP 2026 (Findings)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP.
Terminology
Abstract
Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art embedding models are not reproducible because they rely on closed or undisclosed training data, and they remain insufficiently robust for Southeast Asian languages. We present SEA-LION-Embedding, a fully open and reproducible text-embedding pipeline for Southeast Asian languages trained only on publicly available data, and use it to study three core factors of robust embedding design: data composition, training objective, and base encoder initialization. SEA-LION-Embedding achieves state-of-the-art results on SEA-BED while enabling systematic and reproducible analysis of robust text embeddings for the region.
Sources
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation
- Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
- KaLM-Embedding: Superior Training Data Brings A Stronger Embedding Model
- SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Multilingual E5 Text Embeddings: A Technical Report
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering