Efficient Iterative Retrieval with Heterogeneous Batching
cs.AI, cs.CL
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 15 pages, 8 figures, Accepted to EMNLP 2026 (main conference)
Code: https://github.com/illinoisdata/Orthrus
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Modern information retrieval increasingly employs both embedding and generative models to handle complex queries.
Terminology
Abstract
Modern information retrieval increasingly employs both embedding and generative models to handle complex queries. However, current serving systems suffer from low throughput and poor GPU utilization because they execute these models in isolation. Coarse-grained partitioning, such as dedicating GPUs to specific tasks, fails to adapt to dynamic workloads and creates computational "bubbles". To address these, we present Orthrus, a serving system that performs heterogeneous batching within a unified inference loop. The primary challenge lies in unifying embedding and generation workloads with conflicting computational patterns while optimizing batch composition for high performance. Orthrus addresses these challenges through chunked embedding with incremental pooling and by adjusting batch composition in a workload-aware manner. Evaluation on four A100 GPUs shows that, relative to baseline deployments, Orthrus achieves 1.28 times --4.52 times higher throughput on controlled workloads and up to 55.8% lower end-to-end p99 latency on an iterative-RAG benchmark. We release our code at https://github.com/illinoisdata/Orthrus.
Sources
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
- The Llama 3 Herd of Models
- Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- Unsupervised Dense Information Retrieval with Contrastive Learning
- Cloud-Native Vector Search: A Comprehensive Performance Analysis
- Mistral 7B
- RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
- PHLoRA: data-free Post-hoc Low-Rank Adapter extraction from full-rank checkpoint
- Fast Distributed Inference Serving for Large Language Models
- Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
- Qwen2 Technical Report
- GEM: Empowering LLM for both Embedding Generation and Language Understanding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection