Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression
cs.AI, cs.DC, cs.LG, cs.PF
Submitted: 2026-08-25
Updated: 2026-08-25
Code: https://github.com/Mohammad-Mozaffari/mkor
Project page: https://paramathic.github.io/stoicc-docs
Terminology
Sources
- Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment
- Scalable Second Order Optimization for Deep Learning
- Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers
- 3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMs
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Sparse Networks from Scratch: Faster Training without Losing Performance
- The Llama 3 Herd of Models
- MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
- The State of Sparsity in Deep Neural Networks
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Measuring Massive Multitask Language Understanding
- PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
- Thanos: A Block-wise Pruning Algorithm for Efficient Large Language Model Compression
- Training Recipe for N:M Structured Sparsity with Decaying Pruning Mask
- ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
- MPAX: Mathematical Programming in JAX
- A Practical and Optimal First-Order Method for Large-Scale Convex Quadratic Programming
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection