Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer

summary

Video file (mp4)

The gist

The paper details the development and rigorous evaluation of Arkios, an open bilingual English-Nepali language model trained from scratch.

In short

The episode discusses 'Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.' Hosts examine how this model was built from scratch using 150B tokens of bilingual text. They focus on the importance of script awareness via a custom tokenizer and discuss improvements like manifest-conditioned tool-use contracts for better reliability in low-resource language processing.

Key concepts

Devanagari-Aware Tokenizer
This is a custom tokenizer integrated into the model architecture to correctly handle Nepali script. It solves a structural problem where traditional methods break words at vowel signs, allowing the model to process complex scripts much better than standard methods.
Trained From Scratch
Arkios was built entirely from scratch rather than being fine-tuned on existing models. This involved massive upfront work in data curation and infrastructure setup, using 150B tokens of bilingual English-Nepali text to establish a solid baseline architecture.
Manifest-Conditioned Tool-Use Contract
This improvement restricts the model so it only calls a tool if the user explicitly declares that specific tool is available in the context. This prevents the model from learning to generate arbitrary functions, ensuring its behavior remains predictable and reliable for integration.

Terminology used across episodes

This episode discusses

The paper

Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer · Read on arXiv

Qwen Team

We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single-file C/CUDA training stack and a Devanagari-aware byte-level BPE tokenizer built for this project. On ARC-Easy and ARC-Challenge, Arkios exceeds three comparably sized open models (Pythia-1.4B, TinyLlama-1.1B, OLMo-1B) despite an order of magnitude fewer training tokens, likely aided by a match between our educational-web-text pretraining data and ARC's grade-school-science format rather than a general capability advantage. We report full evaluation results under standard protocols, including a correction to an earlier partial-sample estimate, and findings specific to evaluating small models in a low-resource language: the standard multiple-choice-letter prompt format used by common evaluation harnesses places this model at chance on Nepali reading comprehension, and simultaneously at chance on English in the same format, which would lead a naive benchmark run to conclude the model has no Nepali ability when in fact it does. Concretely, both languages score at chance in the letter-choice format (0.240 Nepali, 0.236 English, against a chance baseline of 0.250), while scoring the answer text directly reveals genuine, English-favoring comprehension (0.306 Nepali, 0.387 English). We describe a manifest-conditioned tool-use contract introduced during instruction tuning, where tool calls are permitted only when a tool manifest is declared in context and suppressed otherwise, and report where that contract holds and where it does not. We release both the base and instruction-tuned model weights under Apache-2.0. The training code and a small privately-sourced portion of the Nepali pretraining corpus are not released; everything needed to reproduce the reported numbers from the released weights is included here.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer".

Tom: The paper details the development and rigorous evaluation of Arkios, an open bilingual English-Nepali language model trained from scratch.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: We started by looking at the title and authors of this paper, "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer," and it immediately tells us where the focus is—a model built entirely from scratch for two specific languages.

Jane: It really highlights that the core innovation isn't just making a large model, but specifically focusing on solving the script problem by integrating a Devanagari-aware tokenizer right into the architecture from the start.

Lu: That focus is key because it means they are tackling a very specific hurdle related to how modern AI systems interact with non-Latin scripts, which is something many researchers overlook when looking at general scaling.

Meng: From an engineering standpoint, starting from scratch on 150B tokens of bilingual text with that custom training stack implies a massive amount of upfront work in data curation and infrastructure setup.

Lalam: That scale shows us how ambitious we can be when we target specific community needs, showing that dedication to a particular language pair can yield significant results.

Tom: And the implications are that if you want high performance in low-resource language processing, you need targeted architectural fixes instead of just hoping a larger model will magically work better.

Jane: Exactly. It suggests that for languages like Nepali, the foundation needs to be built around respecting the character structure, not just throwing more data at it and hoping for the best.

Lu: It points toward a more fundamental way of thinking about multilingual models: treat script awareness as a necessary constraint rather than an optional feature that can be added later.

Meng: That constraint approach is what makes building trustworthy AI systems; you know exactly what structural limitations you're working around before the model even starts generating text.

Lalam: It gives us confidence that we can design systems that genuinely serve communities, not just chase the biggest general benchmarks without considering local linguistic realities.

Tom: So, when we put this together, the title tells us this paper is about a dedicated engineering effort to solve a specific language challenge with a custom tokenizer.

Jane: And it sets the stage perfectly for us to discuss exactly what that means for actual model performance in English and Nepali.

The paper's summary: Tom: We’re moving on now into the summary of "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer" to get a clearer picture of what the model actually does.

Jane: Essentially, the paper explains that Arkios is a 1 point 04B-parameter dense transformer pretrained from scratch using 150B tokens of bilingual English–Nepali text and that it uses this custom Devanagari-aware byte-level BPE tokenizer to handle the script correctly.

Lu: It’s important to understand that the summary emphasizes the scale, which is a 1 point 04B parameter model pretrained on 150B tokens of bilingual text, establishing a solid baseline for what this architecture can do.

Meng: The training details show they used a custom single-file C/CUDA training stack and finished in approximately seventy-nine wall-clock hours on one demand 8xH100 SXM5 node, which is quite a feat for the compute required.

Lalam: That level of scale demonstrates how deep our resources can go when we focus our efforts on a specific task, proving that focused effort can yield substantial results.

Tom: The summary also highlights the architecture details like the layers and hidden size, including attention heads and attention head dimensions FFN dimension SwiGLU, which gives us concrete technical specifications about this model.

Jane: And more importantly for us is what they report on the training data itself: they used a mix of sources like FineWeb-Edu, OpenWebMath GitHub-code-clean, FineWeb-two npi Deva Sangraha (verified Nepali), and the 342B token pool from IndicCorpV2 npi Deva Wikipedia ne Private Nepali data1.

Lu: That specific list of data sources tells us a lot about the diversity of the training material they gathered, showing they were smart about balancing general knowledge with specialized language data.

Meng: That careful selection shows that you have to think very carefully about what kind of pretraining corpus you feed into a model to get the right balance for its intended use.

Lalam: It reminds us that quality and relevance in the training data are just as important as sheer volume when building specialized AI.

Tom: So, we know this model was built from scratch with a specific architecture and a specific set of diverse data sources that we can build on.

Jane: And what they found in their summary is that the model’s capability is heavily tied to the quality and diversity of those 150B tokens of bilingual text.

The paper's improvements: Tom: Now let’s look at the suggested improvements in "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer" to see what they think needs fixing next. The first suggestion is definitely that custom Devanagari-aware byte-level BPE tokenizer.

Jane: This improvement is crucial because it addresses the fundamental way the model sees Nepali text; traditional methods fracture words at vowel signs, which makes Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer capable of handling complex scripts much better.

Lu: It’s a brilliant solution because it addresses what is fundamentally a structural issue with Unicode handling rather than just trying to fix it with more data, making the concept of data-mixture irrelevant for us.

Meng: The C/CUDA training stack is another huge technical improvement; building that custom infrastructure allows for serious hardware efficiency, reaching forty-two point nine five percent utilization on a single H100 node for Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Lalam: It speaks to the idea that we can build highly performant AI systems even when resources are constrained, making advanced AI more accessible globally and improving cultural exchange for us.

Tom: Moving to the second suggestion, they introduce this manifest-conditioned tool-use contract, which means the model only calls a tool if you explicitly declare that specific tool is available in the context of Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Jane: Think of it as giving the model its own internal policy; it refuses to execute a function unless you, as the prompt designer, specifically say that particular tool is available for use right now.

Lu: The masking of that manifest from the training loss is genius because it prevents the model from learning to *generate* those tools itself, keeping its behavior tightly constrained and reliable.

Meng: That constraint ensures predictability in applications; if a real-world software system knows exactly when to expect a tool call from Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer, that simplifies integration immensely.

Lalam: It allows us to build confidence in the AI's decision-making process, knowing it won't invent arbitrary functions or act outside the defined operational boundaries we set for it.

Tom: So these technical fixes are really impressive, but they aren’t all about fixing every single issue in one go; they’re building a more robust system piece by piece.

Jane: That makes sense; you can't fix everything at once, and focusing on the most structurally important components first is a smart strategy for building reliable AI.

Conclusion: Tom: To wrap up our discussion of "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer," we’ve covered the key points about how this model was built and what it actually achieved in terms of performance across different scripts.

Jane: The overall message is that evaluating these models must be more sophisticated, especially when looking at low-resource languages where the standard benchmarks fail to capture true ability for Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Lu: I think the most exciting thing is how this demonstrates that we can achieve high performance by fixing structural issues like tokenization, rather than just trying to scale up indefinitely in the way of Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Meng: It proves that we can build a highly efficient and trustworthy model using clear methods like the tool-use contract and strong hardware utilization for real systems.

Lalam: It gives us hope for all the languages that aren't English by showing what is achievable when we are dedicated to building models like Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer to serve the community.

Tom: We’ve heard some brilliant insights today, but it’s time to wrap up our discussion of Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Jane: Thank you all for joining us on the show and thank you for listening. We'll be back next week with another fascinating paper to discuss.

More episodes

← Home