SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

arXiv:2610.12327 · cs.LG, cs.CL · Submitted 2026-10-08 · Read on arXiv

cs.LG, cs.CL

Submitted: 2026-10-08

Updated: 2026-10-08

Project page: https://wang-qitong.github.io/SparseDecoding

Terminology

Sources

Related papers