Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation
cs.LG, cs.CL
Submitted: 2026-09-02
Updated: 2026-09-02
Code: https://github.com/EleutherAI/lm-evaluation-harness
Terminology
Sources
- PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- On the Geometry of On-Policy Distillation
- Demystifying Low-Rank Knowledge Distillation in Large Language Models: Convergence, Generalization, and Information-Theoretic Guarantees
- Put the Space of LoRA Initialization to the Extreme to Preserve Pre-trained Knowledge
- MiLoRA: Harnessing Minor Singular Components for Parameter-Efficient LLM Finetuning
- The Llama 3 Herd of Models
- Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization
- Distilling the Knowledge in a Neural Network
- Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
- High-Dimensional Search, Low-Dimensional Solution: Decoupling Optimization from Representation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks