Parameter Golf: What Really Works?
cs.CL
Submitted: 2026-07-01
Updated: 2026-09-05
Comments: Accepted at the BabyLM Workshop, EMNLP 2026
Code: https://github.com/PMP56/pmgolf-analysis
License: http://creativecommons.org/licenses/by/4.0/
The gist: How far can a language model improve under a strict artifact budget? Parameter Golf posed this question as an open community challenge in which participants trained the best language model, with the
Terminology
Abstract
How far can a language model improve under a strict artifact budget? Parameter Golf posed this question as an open community challenge in which participants trained the best language model, with the complete artifact (training code + compressed weights) required to fit within 16 MB and to be trained in under ten minutes on 8xH100 SXM GPUs. Quality was measured in bits-per-byte (BPB), the average number of bits required to encode each byte of unseen text. We analyze 2,037 pull requests and 1,430 clean-scored submissions from the contest, build a taxonomy of 84 optimization techniques, and measure each technique's contribution to BPB. The verified leaderboard score dropped from 1.2244 to 1.058 BPB across three phases, a 13.6% reduction, despite individual techniques rarely improving BPB by more than 1%. We show that most techniques' gains shrink when re-measured among competitive submissions, isolating the few methods that help regardless of the surrounding stack. Code and data are available at https://github.com/PMP56/pmgolf-analysis.
Sources
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Improving language models by retrieving from trillions of tokens
- Averaging Weights Leads to Wider Optima and Better Generalization
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
- PaLM: Scaling Language Modeling with Pathways
- Scaling Laws for Neural Language Models
- Generalization through Memorization: Nearest Neighbor Language Models
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Training Compute-Optimal Large Language Models
- SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
- Do Transformer Modifications Transfer Across Implementations and Applications?
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- YaRN: Efficient Context Window Extension of Large Language Models
- LQER: Low-Rank Quantization Error Reconstruction for LLMs
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
- GLU Variants Improve Transformer
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering