ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models
stat.ML, cs.CL, cs.LG
Submitted: 2026-09-03
Updated: 2026-09-03
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Datasheet for the Pile
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Distilling the Knowledge in a Neural Network
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Rethinking Selective Knowledge Distillation
- LLaMA: Open and Efficient Foundation Language Models
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey