ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
cs.OS, cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/Dao-AILab/flash-attention
Terminology
Sources
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- Evaluating Large Language Models Trained on Code
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
- Deep Think with Confidence
- Memory OS of AI Agent
- Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
- Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
- gpt-oss-120b & gpt-oss-20b Model Card
- ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
- Qwen3 Technical Report
- SGLang: Efficient Execution of Structured Language Model Programs
Related papers
- MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices
- Planarian: Managing Agent State with Statepoints
- VUDA: Enabling Controlled Spatial Sharing of Graphics and Compute on NVIDIA GPUs
- GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference
- SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live