Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation
cs.CL, cs.AI, cs.LG
Submitted: 2025-10-25
Updated: 2025-11-07
Comments: Ling 2.0 Technical Report
Code: https://github.com/inclusionAI/Ling-V2
Project page: https://moonshotai.github.io/Kimi-K2
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Program Synthesis with Large Language Models
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- Think you have Solved Direct-Answer Question Answering? Try ARC-DA, the Direct-Answer AI2 Reasoning Challenge
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation
- Evaluating Large Language Models Trained on Code
- Arrows of Math Reasoning Data Synthesis for Large Language Models: Diversity, Complexity and Correctness
- FullStack Bench: Evaluating LLMs as Full Stack Coders
- Training Verifiers to Solve Math Word Problems
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-V3 Technical Report
- Better & Faster Large Language Models via Multi-token Prediction
- InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning
- WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models
- Query-Key Normalization for Transformers
- Training Compute-Optimal Large Language Models
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering