From Zero to Hero: An Open LLM Ecosystem for Armenian
cs.LG, cs.CL
Submitted: 2026-09-03
Updated: 2026-09-03
Comments: 18 pages, 4 figures, 13 tables. Data and model: https://huggingface.co/collections/COPA-AI/armenian-llm-ecosystem. Code: https://github.com/COPATeam/armenian_llm_ecosystem
Code: https://github.com/COPATeam/armenian_llm_ecosystem
License: http://creativecommons.org/licenses/by/4.0/
The gist: Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it.
Terminology
Abstract
Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release two datasets. ArmWeb is an extensively validated corpus of 4.37M Armenian news documents. ArmSTEM is a parallel English-Armenian collection of 373K math and science problems with step-by-step solutions, translated into Armenian and verified through both answer-preserving LLM judgment and human evaluation. Continued pretraining of Gemma-4-E4B on these datasets yields arm-gemma-e4b, which outperforms every existing open Armenian model as well as its unadapted base, and is the first open Armenian LLM with complete training data and recipe. Our ablations show that news-only continued pretraining improves fluency while eroding knowledge, a pattern we also observe in existing Armenian models, and that a small share of verified translated STEM data reverses the loss. We further find that the largest public Armenian corpora overlap web-derived evaluation panels heavily, including a train/test self-overlap inside FineWeb-2. We openly release all data, models, and code.
Sources
- Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models
- Long Chain-of-Thought Reasoning Across Languages
- An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
- Goldfish: Monolingual Language Models for 350 Languages
- Breaking Language Barriers in Multilingual Mathematical Reasoning: Insights and Observations
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Unsupervised Cross-lingual Representation Learning at Scale
- SYSTRAN's Pure Neural Machine Translation Systems
- SambaLingo: Teaching Large Language Models New Languages
- Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
- Training Neural Machine Translation To Apply Terminology Constraints
- FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
- Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
- Latxa: An Open Language Model and Evaluation Suite for Basque
- Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
- Datasheets for Datasets
- SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks