SEA-LION-Embedding: Open and Reproducible Text Embeddings for Southeast Asia

arXiv:2606.03027 · cs.CL · Submitted 2026-06-02 · Read on arXiv

cs.CL

Submitted: 2026-06-02

Updated: 2026-09-09

Comments: Accepted to EMNLP 2026 (Findings)

License: http://creativecommons.org/licenses/by/4.0/

The gist: Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP.

Terminology

Abstract

Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art embedding models are not reproducible because they rely on closed or undisclosed training data, and they remain insufficiently robust for Southeast Asian languages. We present SEA-LION-Embedding, a fully open and reproducible text-embedding pipeline for Southeast Asian languages trained only on publicly available data, and use it to study three core factors of robust embedding design: data composition, training objective, and base encoder initialization. SEA-LION-Embedding achieves state-of-the-art results on SEA-BED while enabling systematic and reproducible analysis of robust text embeddings for the region.

Sources

Related papers