The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness
cs.CL
Submitted: 2026-05-27
Updated: 2026-09-20
Comments: EMNLP 2026 (Main)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property.
Terminology
Abstract
Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding robustness is multidimensional, since models respond differently to different types of variation, and requires dynamic evaluation to expose failures hidden by static benchmarks. We introduce the Harder Text Embedding Benchmark (HTEB), a dynamic evaluation framework that challenges model robustness along three practically interpretable axes (Lexical/Stylistic, Length and Language) by stochastically transforming inputs at evaluation time with an LLM. Evaluating 16 open-weight embedding models on 32 datasets covering 42 languages under transformations validated by 4,800 individual human ratings on an English subsample, supplemented by a Spanish-source evaluation and an exploratory STS-B study of pair-level label preservation, we find three patterns: (1) Models exhibit specific, partly decoupled robustness profiles across axes. (2) Across three model families, scale increases absolute scores but does not close the gap between original and transformed evaluations. Here, scaling tends to improve specifically the Language axis. (3) English datasets are more sensitive to HTEB transformations than multilingual datasets. This demonstrates that HTEB identifies strengths and weaknesses of models along deployment-relevant axes, challenging current embedding benchmarks and arguing for multidimensional, dynamic robustness evaluation. We make the code to run HTEB publicly available.
Sources
- Towards Better Monolingual Japanese Retrievers with Multi-Vector Models
- Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation
- Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
- Gemma 3 Technical Report
- Quantifying Variance in Evaluation Benchmarks
- One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
- Why Comparing Single Performance Scores Does Not Allow to Draw Conclusions About Machine Learning Approaches
- Olmo 3
- Qwen3 Technical Report
- Jasper and Stella: distillation of SOTA embedding models
- Jasper-Token-Compression-600M Technical Report
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World
- Multilingual E5 Text Embeddings: A Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering