Task-Aware Spectral Pruning: A Mixture-of-Masks Framework for Efficient LLM Inference
cs.LG
Submitted: 2026-08-24
Updated: 2026-08-24
Code: https://github.com/tatsu-lab/alpaca_eval
Terminology
Sources
- Program Synthesis with Large Language Models
- Flextron: Many-in-One Flexible Large Language Model
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- DISP-LLM: Dimension-Independent Structural Pruning for Large Language Models
- Diet Your LLM: Dimension-wise Global Pruning of LLMs via Merging Task-specific Importance Score
- LLMs can Compress LLMs: Adaptive Pruning by Agents
- AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
- Rethinking generalization requires revisiting old ideas: statistical mechanics approaches and complex learning behavior
- The Llama 3 Herd of Models
- Certificate-Guided Pruning for Stochastic Lipschitz Optimization
- Frustratingly Easy Task-aware Pruning for Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks