Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference

arXiv:2609.36654 · cs.LG, cs.CL, cs.DC · Submitted 2026-09-29 · Read on arXiv

cs.LG, cs.CL, cs.DC

Submitted: 2026-09-29

Updated: 2026-09-29

Code: https://github.com/vllm-project/llm-compressor

Project page: https://microsoft.github.io/Olive

Terminology

Sources

Related papers