Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights

arXiv:2609.02652 · cs.LG · Submitted 2026-09-02 · Read on arXiv

cs.LG

Submitted: 2026-09-02

Updated: 2026-09-02

Code: https://github.com/ggml-org/llama.cpp

Terminology

Sources

Related papers