CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference
cs.DC, cs.CV, cs.LG
Submitted: 2026-04-07
Updated: 2026-09-15
Comments: 14 pages, 18 figures, 2 tables
Code: https://github.com/lmcache/lmcache
License: http://creativecommons.org/licenses/by/4.0/
The gist: Continuous inference over concurrent video streams imposes substantial compute and memory demands on vision-language model (VLM) serving.
Terminology
Abstract
Continuous inference over concurrent video streams imposes substantial compute and memory demands on vision-language model (VLM) serving. Streaming inference uses sliding windows to maintain a bounded context of recent video, but processing each window independently repeats visual encoding and large language model (LLM) prefilling for similar and overlapping content. Existing optimizations provide limited coordination across these stages and often rely on model-specific training, profiling, or model-generated signals. We present CodecSight, a streaming VLM serving system that uses codec metadata as shared runtime guidance across visual encoding and LLM prefilling, without model-specific training or offline profiling. Codec-derived change signals guide patch pruning before visual encoding, reducing both visual computation and the number of downstream visual tokens. Codec-defined frame types guide selective key-value (KV) refresh across windows, while positional correction enables reuse of the remaining cached keys. Across three VLMs and four video workloads, our vLLM-based implementation supports up to 3.3 times as many concurrent streams and achieves up to a 5.3 times speedup in average time-to-first-token relative to the state-of-the-art baselines. It also reduces executed FLOPs by up to 93%, with a maximum task-quality decrease of 4.64 percentage points.
Sources
- Edge-GPU Based Face Tracking for Face Detection and Recognition Acceleration
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- StreamChat: Chatting with Streaming Video
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- Harnessing Vision-Language Models for Time Series Anomaly Detection
- Stateful Token Reduction for Long-Video Hybrid VLMs
- NoScope: Optimizing Neural Network Queries over Video at Scale
- Process Integrated Computer Vision for Real-Time Failure Prediction in Steel Rolling Mill
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Promptus: Can Prompts Streaming Replace Video Streaming with Stable Diffusion
- MonitorVLM:A Vision Language Framework for Safety Violation Detection in Mining Operations
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
- Efficient Streaming Language Models with Attention Sinks
- PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing