SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor Vasilyev, Min Liu, Weijian Sun, Xin Chen, Yingmeng Gao, Jinhua Zhou, Taolue Chen, Chenwei Wu, Dong Zhang, Wenlong Jin, Jinmin Xiang, Barkova Maria, Ushakov Anton, Xianfei Jin, Tian Ding, Zhihang Lin, Qian Chen, Linxin Yang, Mingzhe Yang, Bingwei Zhang, Hongzhang Yang, Fangxue Zhang, Shijun Qin, Jie Yu, Cuihua Hu, Tolstykh Vasiliy, Nosov Ivan, Abdullin Amir, Zhicheng Zhou, Xin Zhang, Zhixiong Ning, Xutong Zhao, Junjie Huang, Jiajun Liu, Weiyan Kong, Zheng Zhang, Wenhan Luo, Lin Hu, Yangbo Guo, Li Zeng, Shihao Zhang, Baotian Hu, Min Zhang, Haizhou Li, Zhiquan Luo
cs.CL, cs.AI
Submitted: 2026-08-19
Updated: 2026-08-20
Comments: 73 pages, 22 figures, 20 tables
Code: https://github.com/Gurobi/modeling-examples
Project page: https://iiis-ai.github.io/AutoMathText-V2/AutoMathText-V2.pdf
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
- AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units
- Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages
- AVO: Agentic Variation Operators for Autonomous Evolutionary Search
- Training Verifiers to Solve Math Word Problems
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- DeepSeek-V3 Technical Report
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- OR-R1: Automating Modeling and Solving of Operations Research Optimization Problem via Test-Time Reinforcement Learning
- Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
- STARK: Strategic Team of Agents for Refining Kernels
- Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization
- GLM-5: from Vibe Coding to Agentic Engineering
- Measuring Mathematical Problem Solving With the MATH Dataset
- Liger Kernel: Efficient Triton Kernels for LLM Training
- LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages
- Balancing Continuous Pre-Training and Instruction Fine-Tuning: Optimizing Instruction-Following in LLMs
- Kimi K2: Open Agentic Intelligence
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering