Budget-Aware Compression Pipeline for Single-GPU LLM Inference: Methods, Trade-offs, and Coupling Effects
cs.CL
Submitted: 2026-08-30
Updated: 2026-09-03
Terminology
Sources
- SliceGPT: Compress Large Language Models by Deleting Rows and Columns
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- A Survey on Large Language Model Acceleration based on KV Cache Management
- SnapKV: LLM Knows What You are Looking for Before Generation
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
- Pointer Sentinel Mixture Models
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
- ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
- A Simple and Effective Pruning Approach for Large Language Models
- QTIP: Quantization with Trellises and Incoherence Processing
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering