Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
cs.CL, cs.LG
Submitted: 2026-08-27
Updated: 2026-09-03
Comments: 63 pages, 20 figures, 24 tables
Code: https://github.com/thu-pacman/Puro-Megatron
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities.
Terminology
Abstract
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing. Even at a small scale, training Llama-3.2-3B costs over 1.5M, and reproducing SmolLM3-3B needs over 700K. In this report, we present an open pretraining recipe designed to lower this barrier. Using this recipe, we train a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs. The models in the collection differ in token budgets and selected recipe variants. Our best model is trained at a compute cost of less than 6.9K and approaches Qwen2.5-1.5B performance under our evaluation protocol. This cost efficiency is enabled by a combination of approaches, including hardware selection, low-precision training, hyperball optimization, curriculum model averaging, and the data recipe. Beyond the recipe itself, we provide two additional results. First, across the Puro-2B collection, we derive a Puro Cost Scaling Law that relates training cost to average model performance; the fitted law suggests that about 4.4K, less than 5,090, is sufficient to reach the performance of Qwen2-1.5B. Second, as an end-to-end case study, we examine how pretraining data curricula shape downstream performance after post-training. Such controlled studies are enabled by having access to the full pretraining pipeline rather than model weights alone. We release the full training recipe for Puro-2B, including data, code, and model weights under Apache 2.0 at https://huggingface.co/collections/thu-pacman/puro-2b.
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Program Synthesis with Large Language Models
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- Evaluating Large Language Models Trained on Code
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
- DeepSeek-V3 Technical Report
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Gemma 3 Technical Report
- Gemma 4 Technical Report
- WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
- OpenDataLab: Empowering General Artificial Intelligence with Open Datasets
- Scaling Laws for Neural Language Models
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
- LFM2 Technical Report
- Instella: Fully Open Language Models with Stellar Performance
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering