A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
cs.AI
Submitted: 2026-07-03
Updated: 2026-09-07
Code: https://github.com/HarvardMadSys/chutes_workload
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V3 Technical Report
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
- BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
- ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
- WildChat: 1M ChatGPT Interaction Logs in the Wild
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection