Disaggregated Quantization: Specializing LLM Prefill and Decode
cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/unslothai/unsloth
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
- Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
- Progressive Mixed-Precision Decoding for Efficient LLM Inference
- INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
- Adaptive Block-Scaled Data Types
- Extreme Compression of Large Language Models via Additive Quantization
- Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
- Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
- When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference
- MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
- Kimi K3: Open Frontier Intelligence
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- When LLMs get significantly worse: A statistical approach to detect model degradations
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks