An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning
cs.CL, cs.AI, cs.LG
Submitted: 2025-02-17
Updated: 2026-08-29
Comments: Published in Transactions on Machine Learning Research (TMLR), November 2025. Updated to the final published version
Journal ref: Transactions on Machine Learning Research, 2025
Code: https://github.com/CenjhihLi/sparsity_finetuning
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- The Llama 3 Herd of Models
- Zeroth-Order Fine-Tuning of LLMs with Extreme Sparsity
- Scaling Down to Scale Up: A Guide to Parameter-Efficient Fine-Tuning
- Understanding Layer Significance in LLM Alignment
- Embedded Encoder-Decoder in Convolutional Networks Towards Explainable AI
- Gemini: A Family of Highly Capable Multimodal Models
- Neural Network Acceptability Judgments
- A Survey of Resource-efficient LLM and Multimodal Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering