SEA-LION-v4.8: A Technical Report
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-18
Comments: A technical report
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce Nemotron-SEA-LION-v4.8, a family of Southeast Asian Languages In One Network (SEA-LION) models built upon NVIDIA Nemotron 3.
Terminology
Abstract
We introduce Nemotron-SEA-LION-v4.8, a family of Southeast Asian Languages In One Network (SEA-LION) models built upon NVIDIA Nemotron 3. The family includes 30B-A3B and 120B-A12B models, with both continued-pretrained base checkpoints and post-trained variants. We adapt the models using Southeast Asian, reasoning, code, and multilingual parallel data, followed by post-training with supervised fine-tuning and online on-policy distillation. On SEA-HELM, the 30B-A3B model improves the overall SEA score from 46.06 to 51.57, while the 120B-A12B model improves from 49.30 to 63.44. Across seven Southeast Asian languages, we observe broad capability gains with the 120B-A12B model showing broader and more consistent improvements across tasks.
Sources
- Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
- SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding
- Continual Pre-Training of Large Language Models: How to (re)warm your model?
- BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
- Kalahi: A handcrafted, grassroots cultural LLM evaluation suite for Filipino
- OpenSeal: Good, Fast, and Cheap Construction of an Open-Source Southeast Asian LLM via Parallel Data
- Typhoon: Thai Large Language Models
- Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering