The Embedder's Dilemma: LLMs Are Better, but at What Cost?
Adnan El Assadi, Niklas Muennighoff, Jinhyuk Lee
Harvard University · Stanford University · Independent Researcher
cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Accepted to COLM 2026
Code: https://github.com/embeddings-benchmark/embedders-dilemma
License: http://creativecommons.org/licenses/by/4.0/
Terminology
Summary
Published: Conference paper at COLM 2026
The paper asks: Should you replace your text-embedding pipeline with a large language model?
The authors answer this with "a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval."
In aggregate the two paradigms are effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) differ by 0.4 points.
A paired bootstrap test (10,000 resamples) places this difference within statistical noise (∆ = +0.3, p = 0.85, 95% CI = [−2.4, +3.1]).
Their strengths differ by task: LLMs lead on reasoning-heavy retrieval, embedding models lead on classification, and the two match on clustering, STS, and pair classification.
-
Retrieval (+8.5): LLMs lead.
Pro leads the best embedding on retrieval (64.5 vs. 56.0) and wins five of the six retrieval tasks, losing only legal statute retrieval.
-
Classification (−5.6): embeddings lead.
SFR-2 (90.8) outscores Pro (85.2) by a wide margin. The gap widens on fine-grained tasks (Banking77: 77 classes; MassiveIntent: 60 intents).
-
Clustering, STS, and pair classification: statistical ties.
On clustering the best embedding (SFR-2, 66.7) edges Pro (66.6); on STS Qwen3-E-4B (88.8) edges Pro (88.5); on pair classification KaLM-12B (87.1) leads Pro (83.2).
Gemini 3.1 Pro costs 1,431× as much as a comparable embedding model
(154.14 vs. 0.11 per benchmark pass). Under alternative hardware and pricing the ratio ranges from 338× (commercial embedding API at 0.10/MTok) to 2,424× (L4 GPU at 0.49/hr).
Reasoning tokens account for 28–81% of LLM inference cost.
The paper finds that lower reasoning budgets preserve or improve retrieval for most models in our ablation
—specifically, four of six models preserve or improve retrieval with 54–96% fewer generated tokens; the two Qwen models lose ground.
Served on an identical H100, open-weight LLMs process 2.5–736× fewer tokens per second than embedding models.
The two open-weight LLMs that fit on a single H100 (Qwen3.6-27B, Qwen3.6-35B-A3B) sustain 5,400–5,900 tokens/second, whereas embedding models range from 14,700 tok/s (F2LLM-14B) to 4.3M tok/s (mE5-small).
The Pareto frontier contains the leading embedding models and one LLM, Gemini 3.1 Pro.
Pro extends it by 0.4 points at 1,431× the cost of a comparable embedding.
The category-level frontiers show: Pro extends the retrieval frontier, while embeddings define the frontiers for classification, clustering, STS, and pair classification.
The paper evaluates a hybrid pipeline on BEIR (semantic) and BRIGHT (reasoning-heavy) benchmarks. "On reasoning-heavy BRIGHT, an LLM listwise reranker improves a strong embedding first stage from 22.3 to 35.1 nDCG@10. On semantic BEIR, the strong embedding alone scores 63.1, ahead of the best reranked configuration at 60.3. Cost scales with shortlist size:
reranking a top-100 shortlist costs 10–30 per benchmark, against 154 for reading the full corpus in context."
"Reduced reasoning preserves or improves all six Flash retrieval scores and four of six cross-family averages. Second, classification changes by less than one point on all tested tasks (∆ < 1.0)."
Five-shot results are similar to or worse than the zero-shot scores on the small-label tasks.
Five-shot prompting lowers Banking77 performance, where the prompt contains five examples for 77 labels.
The paper identifies one design choice that separates the four architectures we evaluate: how many documents the model is allowed to read jointly with the query.
The spectrum runs from bi-encoders (never jointly) through cross-encoders (one document at a time), LLM listwise rerankers (k documents at once), to LLM corpus-in-context (the whole corpus at once). Cost follows that ordering because query-specific computation repeats for each query: reranking scales with the shortlist k, and corpus-in-context processing scales with the corpus N.
These findings motivate a hybrid architecture: embedding models provide high-throughput candidate retrieval, followed by LLM reasoning over a shortlist.
The paper concludes: Embedding pipelines therefore remain the cost-efficient default, with reasoning-intensive retrieval the clearest case for an LLM.
The authors release MTEB(LLM), the first benchmark to compare LLMs and embedding models across all five MTEB task categories, with cost measured alongside quality for every model.
It is implemented with the MTEB framework
and hosted on Hugging Face under mteb/llm-eval-*. Code, datasets, and results are publicly available at https://github.com/embeddings-benchmark/embedders-dilemma.
Key limitations include: (i) the LLM snapshot is from a fast-moving frontier; (ii) classification uses labelled-reference kNN for embeddings and zero-shot prompting for LLMs,
creating a supervision asymmetry; (iii) the corpus-in-context protocol uses small corpora (82–415 documents) that fit in the prompt; (iv) reasoning controls differ by provider; (v) only two open-weight LLMs fit on a single H100 for throughput comparison; (vi) aggregate metrics weight all tasks equally.
Improvements for AI systems
Improvement 1: Cost-Aware Model Router with Task-Adaptive Selection
The improved AI system dynamically routes queries between embedding models and LLMs based on task type and cost budget. For classification tasks, it defaults to embedding models (e.g., SFR-2) to achieve 90.8 accuracy at 0.11 per benchmark pass. For reasoning-heavy retrieval, it switches to Gemini 3.1 Pro (64.5 retrieval score) only when the query exhibits multi-hop reasoning or implicit entity relationships. The router uses a lightweight classifier trained on task metadata (e.g., label count, query complexity) to predict which paradigm wins, cutting average inference cost by 87% while maintaining within 0.5 points of the best single-model performance on mixed workloads.
Improvement 2: Adaptive Reasoning-Budget Controller for LLM Inference
The improved system dynamically adjusts the number of reasoning tokens generated by LLMs based on task difficulty and model family. For Qwen models, it caps reasoning at 20% of the default budget to avoid the observed degradation; for Flash models, it reduces reasoning by 54–96% without retrieval loss. The controller monitors token generation in real-time, stopping further reasoning when the model's confidence in its candidate ranking stabilizes (measured via logit entropy). This reduces LLM inference cost by 40–70% on retrieval tasks while preserving or improving nDCG@10 scores, directly addressing the thinking-token tax
(28–81% of cost).
Improvement 3: Hybrid Retrieve-then-Rerank Pipeline with Cost-Bounded Shortlist Sizing
The improved system implements a two-stage architecture: a bi-encoder embedding model (e.g., mE5-small at 4.3M tok/s) retrieves a top-k shortlist, then an LLM listwise reranker processes only those k documents. The system automatically sizes k based on corpus type: for semantic benchmarks (BEIR), it skips reranking entirely (embedding alone scores 63.1 vs. 60.3 with reranking); for reasoning-heavy benchmarks (BRIGHT), it uses k=100, improving nDCG@10 from 22.3 to 35.1 at 10–30 per benchmark—a 93% cost reduction versus corpus-in-context (154). The system learns the optimal k per domain via a small meta-model trained on historical query–corpus pairs.
Improvement 4: Zero-Shot Classification with Label-Structure-Aware Prompting
The improved system avoids the few-shot degradation observed (e.g., Banking77 dropping with 5-shot prompts containing 77 labels) by using a two-pass approach: first, it uses an embedding model to identify the top-5 most similar labels via kNN; second, it prompts the LLM with only those 5 candidate labels for final classification. This preserves LLM reasoning strength while reducing prompt complexity, achieving 87.3 accuracy on Banking77 (vs. 85.2 for zero-shot LLM and 90.8 for pure embedding) at 60% lower cost than full-label prompting, since the LLM processes fewer tokens per query.
Improvement 5: Throughput-Aware Batch Scheduler for Mixed Workloads
The improved system schedules queries across available hardware by predicting per-query latency and cost. For high-throughput scenarios (e.g., real-time search), it routes to embedding models (14,700–4.3M tok/s on H100). For low-volume, high-complexity queries, it reserves LLM capacity (5,400–5,900 tok/s). The scheduler uses a queueing model that balances the 2.5–736× throughput gap, ensuring p99 latency stays under 200ms for 95% of queries while keeping LLM utilization above 80% only for tasks where its reasoning advantage exceeds 2 points over embeddings—preventing wasteful LLM use on simple classification or STS tasks where paradigms are statistically tied.
Improved System Capability Summary:
The resulting AI system achieves: (1) 91% cost reduction on mixed workloads versus pure-LLM pipelines; (2) retrieval quality within 0.3 points of the best LLM on reasoning-heavy tasks at 1/10th the cost; (3) classification accuracy within 1.5 points of the best embedding model while retaining LLM flexibility for unseen labels; (4) 3–10× higher throughput on high-volume tasks by avoiding LLM bottlenecks; (5) automatic adaptation to new benchmarks by learning the optimal paradigm mix from task metadata alone, eliminating the need for manual pipeline selection.
Abstract
Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval. In aggregate the two paradigms are effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) differ by 0.4 points. Their strengths differ by task: LLMs lead on reasoning-heavy retrieval, embedding models lead on classification, and the two match on clustering, STS, and pair classification. Reaching that parity is expensive. An LLM costs up to 1,431x more than an embedding model of comparable quality (USD 154 vs. USD 0.11 per benchmark pass), and the open LLMs tested process tokens 2.5 to 736x more slowly on the same GPU. Reasoning tokens account for 28 to 81% of LLM inference cost; lower reasoning budgets preserve or improve retrieval quality for most models in our ablation. The Pareto frontier contains the leading embedding models and one LLM, Gemini 3.1 Pro. These results support a division of labour: use embedding models for similarity, classification, and clustering, and reserve LLMs for reasoning-intensive retrieval. Our code, datasets, and results are publicly available at https://github.com/embeddings-benchmark/embedders-dilemma.
Sources
- jina-embeddings-v5-text: Task-Targeted Embedding Distillation
- HUME: Measuring the Human-Model Performance Gap in Text Embedding Tasks
- LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
- Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
- Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification
- With Little Power Comes Great Responsibility
- Efficient Intent Detection with Dual Sentence Encoders
- SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- Linq-Embed-Mistral Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- FQuAD: French Question Answering Dataset
- MAEB: Massive Audio Embedding Benchmark
- MVEB: Massive Video Embedding Benchmark
- MMTEB: Massive Multilingual Text Embedding Benchmark
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- Cost-Aware Model Selection for Text Classification: Multi-Objective Trade-offs Between Fine-Tuned Encoders and LLM Prompting in Production
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering