On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/QwenLM/FlashQLA
Terminology
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm
- Program Synthesis with Large Language Models
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
- Alternating Updates for Efficient Transformers
- Language Models are Few-Shot Learners
- Evaluating Large Language Models Trained on Code
- Beyond N-gram: Data-Aware X-GRAM Extraction for Efficient Embedding Parameter Scaling
- PaLM: Scaling Language Modeling with Pathways
- Training Verifiers to Solve Math Word Problems
- Scaling Vision Transformers to 22 Billion Parameters
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- Practical Efficiency of Muon for Pretraining
- SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
- Are We Done with MMLU?
- Gemma 4 Technical Report
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Deep Residual Learning for Image Recognition
- Measuring Mathematical Problem Solving With the MATH Dataset
- RULER: What's the Real Context Size of Your Long-Context Language Models?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering