Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference
cs.LG, cs.CL, cs.DC
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/vllm-project/llm-compressor
Project page: https://microsoft.github.io/Olive
Terminology
Sources
- SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Low-Rank Quantization-Aware Training for LLMs
- Evaluating Large Language Models Trained on Code
- EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
- Training Verifiers to Solve Math Word Problems
- Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling
- Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- The Llama 3 Herd of Models
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to Giant
- ScaleSweep: Accurate NVFP4 Post-Training Quantization of LLMs via Block Scale Initialization
- LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
- SpinQuant: LLM quantization with learned rotations
- ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
- Pretraining Large Language Models with NVFP4
- TorchAO: PyTorch-Native Training-to-Serving Model Optimization
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Microscaling Data Formats for Deep Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks