LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum
Zhichao Xu, Shengyao Zhuang, Crystina Zhang, Xueguang Ma, Yijun Tian, Maitrey Mehta, Jimmy Lin, Vivek Srikumar
cs.IR, cs.CL
Submitted: 2026-08-19
Updated: 2026-08-20
Comments: SIGIR 2026 camera ready
License: http://creativecommons.org/licenses/by/4.0/
The gist: While dense retrieval models have been the standard for state-of-the-art information retrieval, their deployment is often constrained by high memory requirements and reliance on GPU accelerators for
Terminology
Abstract
While dense retrieval models have been the standard for state-of-the-art information retrieval, their deployment is often constrained by high memory requirements and reliance on GPU accelerators for vector similarity search at scale. Learned sparse retrieval offers a compelling alternative by enabling efficient search via inverted indices, yet it has historically received less attention than dense approaches. In this paper, we introduce LACONIC, a family of learned sparse retrievers based on the Llama3 architecture (1B, 3B, and 8B). We propose a streamlined two-phase training curriculum consisting of (1) weakly supervised pre-finetuning to adapt causal LLMs for bidirectional contextualization and (2) high-signal finetuning using curated hard negatives. Our results demonstrate that LACONIC effectively bridges the performance gap with dense models: the 8B variant achieves a state-of-the-art 60.2 nDCG@10 on the MTEB Retrieval benchmark, ranking 15th on the leaderboard as of February 5th, 2026, while utilizing 74% less index memory than an equivalent dense model. By delivering high retrieval effectiveness on commodity CPU hardware with a fraction of the compute budget required by competing models, LACONIC provides a scalable and efficient solution for real-world search applications. We fully open source our code implementation and trained checkpoints to facilitate reproducibility.
Sources
- SparTerm: Learning Term-based Sparse Representation for Fast Text Retrieval
- Efficient Sketching and Nearest Neighbor Search Algorithms for Sparse Vector Sets
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Luxical: High-Speed Lexical-Dense Text Embeddings
- Mistral-SPLADE: LLMs for better Learned Sparse Retrieval
- SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval
- Tevatron: An Efficient and Flexible Toolkit for Dense Retrieval
- Towards Competitive Search Relevance For Inference-Free Learned Sparse Retrievers
- The Llama 3 Herd of Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Billion-scale similarity search with GPUs
- SPLADE-v3: New baselines for SPLADE
- Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector
- Nomic Embed: Training a Reproducible Long Context Text Embedder
- Representation Learning with Contrastive Predictive Coding
- Repetition Improves Language Model Embeddings
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- RankMamba: Benchmarking Mamba's Document Ranking Performance in the Era of Transformers
- CSPLADE: Learned Sparse Retrieval with Causal Language Models
- Distillation versus Contrastive Learning: How to Train Your Rerankers
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG