K2-V2: A 360-Open, Reasoning-Enhanced LLM
cs.LG
Submitted: 2025-12-05
Updated: 2026-09-17
Code: https://github.com/jgm/pandoc
Project page: https://moonshotai.github.io/Kimi-K2/thinking.html
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Nemotron-4 340B Technical Report
- gpt-oss-120b & gpt-oss-20b Model Card
- Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
- Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
- SantaCoder: don't reach for the stars!
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Program Synthesis with Large Language Models
- Qwen Technical Report
- Efficient Training of Language Models to Fill in the Middle
- Stable LM 2 1.6B Technical Report
- Llama-Nemotron: Efficient Reasoning Models
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- Evaluating Large Language Models Trained on Code
- K2-Think: A Parameter-Efficient Reasoning System
- Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks