ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
cs.AI
Submitted: 2026-09-11
Updated: 2026-09-20
Code: https://github.com/zgcagi/ZGCM-1
License: http://creativecommons.org/licenses/by/4.0/
The gist: In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency.
Terminology
Abstract
In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use. To support this paradigm across a 256K context, we develop an end-to-end, high-efficiency open training recipe: Architecture & System Co-design: interleaved gated sliding-window and full attention, and a stable FP8 Muon optimizer; Progressive Curriculum & MDP Mid-Training: context scaling across 16K, 64K, and 256K, and the reformulation of interaction traces into Markov Decision Processes. Furthermore, we establish an AI-native R&D workflow where agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation. Extensive evaluations show that ZGCM-1-7B is competitive across 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1. We also show that our pre-training design offers a 4.2x efficiency improvement in 16K pre-training time-to-loss. Across the full development lifecycle, we distill eight actionable empirical findings-spanning architectural scaling, SFT quality pruning, long-context generalization, and agentic co-training dynamics. To facilitate community research, we open-source model weights from the pre-training, mid-training, and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs.
Sources
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Llemma: An Open Language Model For Mathematics
- Longformer: The Long-Document Transformer
- Generating Long Sequences with Sparse Transformers
- On the Measure of Intelligence
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Scaling Vision Transformers to 22 Billion Parameters
- HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
- Gemma 3 Technical Report
- GLM-5: from Vibe Coding to Agentic Engineering
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Kimi K3: Open Frontier Intelligence
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Deduplicating Training Data Makes Language Models Better
- TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection