EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
summary
The gist
The paper addresses the critical need to enhance the capabilities of large multilingual language models (LLMs) specifically for Estonian, a language that often lacks sufficient digital resources
In short
The episode discusses a paper titled "EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training." Hosts examine how this methodology uses Continued Pretraining (CPT) on a large, diverse dataset to improve the model's performance in Estonian. They conclude that these adaptation strategies are crucial for achieving language equity and cultural representation.
Key concepts
- EstLLM
- EstLLM is a large language model designed to enhance Estonian capabilities. It is built upon the Llama 3.1 8B base model, which was then adapted using Continued Pretraining (CPT) on a curated dataset to make the model highly proficient in Estonian.
- Continued Pretraining (CPT)
- This is a methodology used in the paper where a base LLM is given targeted exposure to Estonian language data. It involves training the model on specific, high-quality text to teach it nuances and improve its linguistic competence without losing general knowledge.
- Chat Vector Merging
- This technique was used after CPT to address instruction-following issues. It allows researchers to inject alignment properties from an instruction-tuned version back into the CPT model, helping the model follow commands while retaining its new Estonian knowledge.
Terminology used across episodes
This episode discusses
- EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training · Paper Radio
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Apertus: Democratizing Open and Compliant LLMs for Global Language Environments
- Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
- On Tiny Episodic Memories in Continual Learning
- To Code, or Not To Code? Exploring Impact of Code in Pre-training
- Instruction Pre-Training: Language Models are Supervised Multitask Learners
- Training Verifiers to Solve Math Word Problems
- Gemma 3 Technical Report
- Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
- The Llama 3 Herd of Models · Paper Radio
- Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
- Textbooks Are All You Need
- MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
- Simple and Scalable Strategies to Continually Pre-train Large Language Models
- EMMA-500: Enhancing Massively Multilingual Adaptation of Large Language Models
- DataComp-LM: In search of the next generation of training sets for language models
- FastText.zip: Compressing text classification models
- Bag of Tricks for Efficient Text Classification
- Estonian Native Large Language Model Benchmark
- No Language Left Behind: Scaling Human-Centered Machine Translation
The paper
EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training · Read on arXiv
Aleksei Dorkin, Taido Purason, Emil Kalbaliyev, Hele-Andra Kuulmets, Marii Ojastu, Mark Fišel, Tanel Alumäe, Eleri Aedmaa, Krister Kruusmaa, Kairit Sirts
University of Tartu · Tallinn University of Technology · Institute of the Estonian Language · Tallinn University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training".
Jane: The paper was written by Aleksei Dorkin, Taido Purason, Emil Kalbaliyev, Hele-Andra Kuulmets, Marii Ojastu et al. from University of Tartu (Institute of Computer Science) and Tallinn University of Technology (Department of Software Science) and Institute of the Estonian Language and Tallinn University (School of Humanities).
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary: Tom: The paper outlines a very specific methodology in its summary, and it’s not just one step. They are taking Llama three point one 8B as their base model, which is a large, powerful LLM that already has some multilingual abilities.
Jane: But the core idea is that this base model needs more targeted exposure to Estonian, so they apply Continued Pretraining or CPT on a carefully constructed data mixture designed to teach it the language.
Lu: I’m fascinated by the composition of that mix—it's not just random Estonian text; it includes English replay and code and math, which is a very sophisticated way to ensure they don't lose its general knowledge.
Meng: From an engineering standpoint, that thirty-five point seven billion token mixture is quite impressive to manage; it shows how much data they needed to shift the model’s focus onto Estonian without just overwhelming it with sheer volume.
Lalam: It's about balancing that exposure—the blend of Estonian and general knowledge—to create a truly capable language model that respects the nuances of the language.
Tom: That mixture, containing eight point six billion tokens from the Estonian National Corpus, really forms the foundation of their work on "EstLLM."
Jane: It’s a very deliberate process, but we need to look at how this continued pretraining leads to actual changes in performance before we move onto the results.
Improvements: Tom: The paper's core findings demonstrate clear and consistent gains across various benchmarks for "EstLLM." They aren't just getting better at Estonian; they’re showing improvement in linguistic competence and knowledge too.
Jane: It seems the CPT process really builds a stronger foundation, but it also suggests that after continued pretraining, we need more specific tuning to make the model follow instructions well.
Lu: I was particularly interested in their use of "chat vector merging" to address this; it’s such an elegant way to inject the alignment properties from an instruction-tuned version back into the CPT model without having to retrain everything.
Meng: That technique, combining the base model with a chat vector difference, is very clever because it lets us recover instruction-following abilities while keeping that Estonian knowledge we just gained.
Lalam: The improvements in reasoning and translation are also impressive; they’re not just memorizing Estonian words, they’ are applying complex logic in the language.
Tom: And Jane mentioned that this isn's a perfect trade-off, right? The authors show that while EstLLM improves dramatically on Estonian tasks, it maintains competitive performance on English benchmarks.
Jane: Yes, but the fact that the results are consistent across automatic and comparative evaluations suggests a robust model without sacrificing its general capabilities.
Tom: That's a powerful finding; when we see improvements in one language without degradation in another area of achieving proficiency, it' something we should all be proud of "EstLLM."
Jane: It’s an encouraging result that leads us straight into the final wrap-up where we can discuss what this all means for global AI.
Conclusion: Tom: So, we have seen the journey of "EstLLM," from the initial problem of English dominance to how they solved it using CPT and post-training methods. The evidence is pretty compelling that these adaptation strategies work.
Jane: It's more than just a technical success for this specific project; it’ a roadmap for improving language equity in AI, which is a huge win for the whole team.
Lu: I think the implications are massive; if "EstLLM" proves successful, we are setting precedents for how diverse, smaller languages will be integrated into future global models.
Meng: Practically speaking, this demonstrates that we don're finding computationally inexpensive ways to achieve high performance after language adaptation without having to start from scratch every single time.
Lalam: This is about making sure the voices of all cultures are heard; "EstLLM" offers a path toward a more globally representative and culturally sensitive AI landscape for everyone.
Tom: I think we have a lot to process, but we want to wrap up this discussion on "EstLLM" by sharing our final thoughts with the listeners.
Jane: It's an exciting time for language technology, and it’s clear that this work is pushing the boundaries of what is possible.
Lu: I hope this research opens up possibilities for more diverse creative expressions in Estonian through AI too, which has immense potential for cultural growth.
Meng: I just hope we can see these methods applied to make a practical difference in deployment across different regions now.
Lalam: We need to keep the momentum going, ensuring that this success is the starting point, not the finish line for a broader language adaptation movement.
Tom: That's a great way to end our discussion on "EstLLM." Thank you all for helping us explore these concepts today!
Conclusion: Tom: So, we've covered so much ground today talking through "EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training." It really shows how crucial localized data is for making these big models actually useful.
Jane: Exactly, Tom. What I take away from all of this is that it’s not enough just to have a massive model; you genuinely need to tailor it to the specific nuances of a language community, like Estonian.
Lu: That idea of tailoring really unlocks something incredible; we're talking about making these systems culturally aware, not just grammatically correct.
Meng: But practically speaking, Jane pointed out that local data is key—I wonder how much compute time it takes to successfully integrate a localized pretraining phase like this into an existing infrastructure?
Jane: Well, Meng, I think the breakthrough here isn't just the compute; it’s proving that continued pretraining *works* effectively enough to move the needle significantly.
Tom: And Lu was right about the cultural awareness—it's not just translation; it’s understanding context and local idioms, which is a huge leap for general AI models.
Lu: It suggests a paradigm shift where we view language models less as universal engines and more as deeply rooted digital extensions of specific cultures.
Meng: If that’s true, then the next hurdle for me is making the deployment scalable across dozens of small, specialized languages worldwide, not just one region.
Lalam: What Meng brings up about scaling is profound because it touches on how knowledge itself flows; if we can master fine-tuning for niche language groups, we elevate global human culture by giving voice to every dialect.
Jane: It makes you think about what happens when these tools become commonplace, doesn't it? They could genuinely democratize access to high-quality AI assistance everywhere.
Tom: Right, because the implications are huge—it means that smaller linguistic communities aren't left behind just because they don't speak English or Mandarin.
Lu: Looking forward, this work paves the way for hyper-localized digital assistants that can assist in everything from historical preservation to modern scientific fields within Estonia.
Meng: So, instead of just viewing it as a single model upgrade, we should see this methodology as a toolkit that any regional tech hub can adopt right away.
Lalam: Ultimately, the most impactful vision is seeing these enhanced models foster greater cross-cultural understanding by serving as bridges between language groups.
Jane: It's been such an exciting discussion wrapping up our look at "EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training."
Tom: Thanks to all of you for breaking this down so well; we really appreciate the insights from Lu, Meng, and Lalam today.
Lu: It's been a wild ride through the possibilities!
Meng: Great discussion; I feel much clearer on the engineering path forward now.
Lalam: Remember that this advancement helps weave a richer tapestry of global human expression.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language