Celty: SpMSpV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
cs.AR, cs.LG
Submitted: 2026-08-02
Updated: 2026-09-05
Comments: ICCAD 2026. The code is available on Github at https://github.com/RuokaiYin/Celty
Code: https://github.com/RuokaiYin/Celty
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security
- MACKO: Sparse Matrix-Vector Multiplication for Low Sparsity
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
- Pointer Sentinel Mixture Models
- 2SSP: A Two-Stage Framework for Structured Pruning of LLMs
- A Simple and Effective Pruning Approach for Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
- DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
- OPT: Open Pre-trained Transformer Language Models
- gpt-oss-120b & gpt-oss-20b Model Card
- Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition
- BlockPruner: Fine-grained Pruning for Large Language Models
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4