The Embedder's Dilemma: LLMs Are Better, but at What Cost?
AuthorsAdnan El Assadi, Niklas Muennighoff, Jinhyuk Lee
LLMs can match embedding models on text tasks, but dedicated embedders usually deliver similar quality at a tiny fraction of the cost.
Key results
Gemini 3.1 Pro scored 77.6 versus Octen-8B at 77.2 on MTEB(LLM).
Embedding models led LLMs by 5.6 points; SFR-2 scored 90.8 versus Gemini 3.1 Pro at 85.2.
LLMs led embeddings by 8.5 points, with Gemini 3.1 Pro at 64.5 versus Octen-8B at 56.0.
A Gemini 3.1 Pro benchmark pass cost $154 versus $0.11 for Octen-8B.
On the same NVIDIA H100, embedding models processed up to 736 times more tokens per second than open LLMs.
What the paper found
The paper introduces MTEB(LLM), a controlled cost-aware comparison of 10 LLMs and 26 embedding models across 37 tasks covering classification, semantic textual similarity, clustering, pair classification, and retrieval. Google’s Gemini 3.1 Pro scored 77.6 overall, while the embedding model Octen-8B reached 77.2, leaving only a 0.4-point difference. The task breakdown is more important than the aggregate: embedding models led classification by 5.6 points, with SFR-2 scoring 90.8 versus Gemini’s 85.2, while LLMs led reasoning-heavy retrieval by 8.5 points, scoring 64.5 versus Octen-8B’s 56.0. Clustering, STS, and pair classification were effectively tied. That parity is expensive: one Gemini 3.1 Pro benchmark pass cost $154 versus $0.11 for Octen-8B, a 1431-times gap. On the same NVIDIA H100, open LLMs processed 2.5 to 736 times fewer tokens per second than embedding models. Reasoning tokens consumed 28 to 81 percent of LLM inference cost, yet reducing reasoning tokens by 54 to 96 percent preserved or improved retrieval for most models. On BRIGHT, an LLM reranker improved Qwen3-E-8B retrieval from 22.3 to 35.1 nDCG@10, whereas on semantic BEIR, the embedding model alone scored 63.1, beating reranked variants. The recommendation is a hybrid pipeline: embeddings for high-throughput similarity, classification, and clustering, with LLMs such as Gemini, DeepSeek-R1, or Qwen reserved for reasoning-intensive retrieval; OpenAI’s GPT-5 and Anthropic’s Claude Opus 4.6 were discussed but not evaluated.
Original abstract
Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval. In aggregate the two paradigms are effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) differ by 0.4 points. Their strengths differ by task: LLMs lead on reasoning-heavy retrieval, embedding models lead on classification, and the two match on clustering, STS, and pair classification. Reaching that parity is expensive. An LLM costs up to 1,431x more than an embedding model of comparable quality (USD 154 vs. USD 0.11 per benchmark pass), and the open LLMs tested process tokens 2.5 to 736x more slowly on the same GPU. Reasoning tokens account for 28 to 81% of LLM inference cost; lower reasoning budgets preserve or improve retrieval quality for most models in our ablation. The Pareto frontier contains the leading embedding models and one LLM, Gemini 3.1 Pro. These results support a division of labour: use embedding models for similarity, classification, and clustering, and reserve LLMs for reasoning-intensive retrieval. Our code, datasets, and results are publicly available at https://github.com/embeddings-benchmark/embedders-dilemma.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.