ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
cs.LG, cs.CL
Submitted: 2025-05-26
Updated: 2026-09-21
Comments: published in Transactions on Machine Learning Research (TMLR)
Code: https://github.com/karpathy/nanoGPT
Project page: http://skylion007.github.io/OpenWebTextCorpus
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model pretraining is compute-intensive, yet many tokens contribute marginally to learning, resulting in inefficiency.
Terminology
Abstract
Large language model pretraining is compute-intensive, yet many tokens contribute marginally to learning, resulting in inefficiency. We introduce Efficient Selective Language Modeling (ESLM), a risk-aware algorithm that improves training efficiency and distributional robustness by performing online token-level batch selection. ESLM leverages per-token statistics (e.g., entropy or loss) and applies value-at-risk thresholding to retain only the most informative tokens per batch. This data-centric mechanism reshapes the training loss, prioritizing high-risk tokens and eliminating redundant gradient computation. We frame ESLM as a bilevel game: the model competes with a masking adversary that selects worst-case token subsets under a constrained thresholding rule. In the loss-based setting, ESLM recovers conditional value-at-risk loss minimization, providing a principled connection to distributionally robust optimization. We extend our approach to Ada-ESLM, which adaptively tunes the selection confidence during training. Experiments on GPT-2 pretraining show that ESLM significantly reduces training FLOPs while maintaining or improving both perplexity and downstream performance compared to baselines. Our approach also scales across model sizes, pretraining corpora, and integrates naturally with knowledge distillation.
Sources
- A Survey on Data Selection for Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Irreducible Curriculum for Language Model Pretraining
- DoGE: Domain Reweighting with Generalization Estimation
- Distilling the Knowledge in a Neural Network
- Accelerating Deep Learning by Focusing on the Biggest Losers
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Scaling Laws for Neural Language Models
- Distributionally Robust Optimization
- Rho-1: Not All Tokens Are What You Need
- Online Batch Selection for Faster Training of Neural Networks
- When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale
- LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
- Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
- Distributionally Robust Language Modeling
- A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
- How to Train Data-Efficient LLMs
- Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining
- Crowdsourcing Multiple Choice Science Questions
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks