ufakzeka-1: Building and Evaluating a 151M-Parameter Turkish Language Model from Scratch
cs.CL
Submitted: 2026-09-18
Updated: 2026-09-18
Comments: 8 pages. Model, code, evaluation results and demo: https://ufakzeka.com, https://huggingface.co/ufakai, https://github.com/ufakai/ufakzeka
Code: https://github.com/ufakai/ufakzeka
License: http://creativecommons.org/licenses/by/4.0/
The gist: We describe ufakzeka-1, a 151M-parameter (182M with embeddings) decoder-only Turkish language model pretrained from scratch on 13.5B tokens of openly licensed text and instruction-tuned for chat, at
Terminology
Abstract
We describe ufakzeka-1, a 151M-parameter (182M with embeddings) decoder-only Turkish language model pretrained from scratch on 13.5B tokens of openly licensed text and instruction-tuned for chat, at a total cost of about 286 in cloud GPU, API and notebook time. The contribution is not the model's capability, which is what a model this size can be expected to have, but the record of building and measuring it: a Turkish byte-level tokenizer at 1.77 tokens per word, a three-stage pretraining schedule, a post-training mixture of openly licensed and generated data, and an evaluation battery of release gates, a rule-checked sweep of 5,508 conversations, judged conversations and hand tests, all with prompts held out from the training data, enforced by decontamination inside the data build and by a checked-in invariant script we run before each build. We report three findings that we believe transfer to other small-model efforts: a safety gate that had been "fixed" with training data written from its own questions read 64/64 while the honest figure was 34/64; training-seed variance was as large as the spread across every recipe we tried, so single-seed comparisons at this scale are uninformative; and data rounds repaired only what was absent from the data, while identity tracking over long context and multi-turn arithmetic did not move across any data change we tried, which we read as limits of the model size rather than gaps in the data, a reading the next, larger model will test. Weights, the data recipe, the evaluation code and the spend ledger are released under Apache-2.0.
Sources
- Qwen3 Technical Report
- Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
- LLaMA: Open and Efficient Foundation Language Models
- PaLM: Scaling Language Modeling with Pathways
- Number Cookbook: Number Understanding of Language Models and How to Improve It
- Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
- Muon is Scalable for LLM Training
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- On the Effect of Instruction Tuning Loss on Generalization
- Replaying pre-training data improves fine-tuning
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning
- Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models
- HellaSwag: Can a Machine Really Finish Your Sentence?
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering