Made in Hungary: Comments on the performance of generative language models
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 14 pages, 1 figure
Code: https://github.com/nytud/HuLU
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: In recent years, three initiatives have emerged to develop generative language models in Hungary.
Terminology
Abstract
In recent years, three initiatives have emerged to develop generative language models in Hungary. The motivation behind them is the same. For Hungarian, no model with the given capability existed, or existing English-centric models offered limited proficiency. A detailed examination of the corresponding studies, however, reveals several methodological limitations. First, the reliability of the evaluation protocols is questionable. Contrary to the findings of Csibi et al. [2026], evaluation under the recommended inference settings shows that Qwen3-4B achieves higher scores than Racka-4B, its Hungarian-adapted version. Data contamination is evident in the work of Yang et al. [2025d] and Szentmihályi et al. [2025], potentially biasing the reported results. Second, the training pipelines fall short of current best practices in corpus curation and data mixture, which risks wasting substantial compute on low-quality data. The lack of controlled ablations prevents reliable assessment of these choices. Third, none of the three papers assessed forgetting or capability loss. Testing the adapted models on a subset of the original benchmarks indicates performance decline in all three cases, especially Racka-4B. These observations emphasize the importance of rigorous experimental design in language model development, given the significant computational and financial costs involved.
Sources
- Phi-4 Technical Report
- Flamingo: a Visual Language Model for Few-Shot Learning
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Stable LM 2 1.6B Technical Report
- LoRA Learns Less and Forgets Less
- AudioLM: a Language Modeling Approach to Audio Generation
- Language Models are Few-Shot Learners
- Evaluating Large Language Models Trained on Code
- PaLM: Scaling Language Modeling with Pathways
- Training Verifiers to Solve Math Word Problems
- SambaLingo: Teaching Large Language Models New Languages
- Efficiently Adapting Pretrained Language Models To New Languages
- Racka: Efficient Hungarian LLM Adaptation on Academic Infrastructure
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Salamandra Technical Report
- The Llama 3 Herd of Models
- Training Compute-Optimal Large Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering