Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer

arXiv:2608.30092 · cs.CL, cs.AI · Submitted 2026-08-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer".

Tom: The paper details the development and rigorous evaluation of Arkios, an open bilingual English-Nepali language model trained from scratch.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: We started by looking at the title and authors of this paper, "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer," and it immediately tells us where the focus is—a model built entirely from scratch for two specific languages.

Jane: It really highlights that the core innovation isn't just making a large model, but specifically focusing on solving the script problem by integrating a Devanagari-aware tokenizer right into the architecture from the start.

Lu: That focus is key because it means they are tackling a very specific hurdle related to how modern AI systems interact with non-Latin scripts, which is something many researchers overlook when looking at general scaling.

Meng: From an engineering standpoint, starting from scratch on 150B tokens of bilingual text with that custom training stack implies a massive amount of upfront work in data curation and infrastructure setup.

Lalam: That scale shows us how ambitious we can be when we target specific community needs, showing that dedication to a particular language pair can yield significant results.

Tom: And the implications are that if you want high performance in low-resource language processing, you need targeted architectural fixes instead of just hoping a larger model will magically work better.

Jane: Exactly. It suggests that for languages like Nepali, the foundation needs to be built around respecting the character structure, not just throwing more data at it and hoping for the best.

Lu: It points toward a more fundamental way of thinking about multilingual models: treat script awareness as a necessary constraint rather than an optional feature that can be added later.

Meng: That constraint approach is what makes building trustworthy AI systems; you know exactly what structural limitations you're working around before the model even starts generating text.

Lalam: It gives us confidence that we can design systems that genuinely serve communities, not just chase the biggest general benchmarks without considering local linguistic realities.

Tom: So, when we put this together, the title tells us this paper is about a dedicated engineering effort to solve a specific language challenge with a custom tokenizer.

Jane: And it sets the stage perfectly for us to discuss exactly what that means for actual model performance in English and Nepali.

The paper's summary: Tom: We’re moving on now into the summary of "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer" to get a clearer picture of what the model actually does.

Jane: Essentially, the paper explains that Arkios is a 1 point 04B-parameter dense transformer pretrained from scratch using 150B tokens of bilingual English–Nepali text and that it uses this custom Devanagari-aware byte-level BPE tokenizer to handle the script correctly.

Lu: It’s important to understand that the summary emphasizes the scale, which is a 1 point 04B parameter model pretrained on 150B tokens of bilingual text, establishing a solid baseline for what this architecture can do.

Meng: The training details show they used a custom single-file C/CUDA training stack and finished in approximately seventy-nine wall-clock hours on one demand 8xH100 SXM5 node, which is quite a feat for the compute required.

Lalam: That level of scale demonstrates how deep our resources can go when we focus our efforts on a specific task, proving that focused effort can yield substantial results.

Tom: The summary also highlights the architecture details like the layers and hidden size, including attention heads and attention head dimensions FFN dimension SwiGLU, which gives us concrete technical specifications about this model.

Jane: And more importantly for us is what they report on the training data itself: they used a mix of sources like FineWeb-Edu, OpenWebMath GitHub-code-clean, FineWeb-two npi Deva Sangraha (verified Nepali), and the 342B token pool from IndicCorpV2 npi Deva Wikipedia ne Private Nepali data1.

Lu: That specific list of data sources tells us a lot about the diversity of the training material they gathered, showing they were smart about balancing general knowledge with specialized language data.

Meng: That careful selection shows that you have to think very carefully about what kind of pretraining corpus you feed into a model to get the right balance for its intended use.

Lalam: It reminds us that quality and relevance in the training data are just as important as sheer volume when building specialized AI.

Tom: So, we know this model was built from scratch with a specific architecture and a specific set of diverse data sources that we can build on.

Jane: And what they found in their summary is that the model’s capability is heavily tied to the quality and diversity of those 150B tokens of bilingual text.

The paper's improvements: Tom: Now let’s look at the suggested improvements in "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer" to see what they think needs fixing next. The first suggestion is definitely that custom Devanagari-aware byte-level BPE tokenizer.

Jane: This improvement is crucial because it addresses the fundamental way the model sees Nepali text; traditional methods fracture words at vowel signs, which makes Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer capable of handling complex scripts much better.

Lu: It’s a brilliant solution because it addresses what is fundamentally a structural issue with Unicode handling rather than just trying to fix it with more data, making the concept of data-mixture irrelevant for us.

Meng: The C/CUDA training stack is another huge technical improvement; building that custom infrastructure allows for serious hardware efficiency, reaching forty-two point nine five percent utilization on a single H100 node for Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Lalam: It speaks to the idea that we can build highly performant AI systems even when resources are constrained, making advanced AI more accessible globally and improving cultural exchange for us.

Tom: Moving to the second suggestion, they introduce this manifest-conditioned tool-use contract, which means the model only calls a tool if you explicitly declare that specific tool is available in the context of Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Jane: Think of it as giving the model its own internal policy; it refuses to execute a function unless you, as the prompt designer, specifically say that particular tool is available for use right now.

Lu: The masking of that manifest from the training loss is genius because it prevents the model from learning to *generate* those tools itself, keeping its behavior tightly constrained and reliable.

Meng: That constraint ensures predictability in applications; if a real-world software system knows exactly when to expect a tool call from Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer, that simplifies integration immensely.

Lalam: It allows us to build confidence in the AI's decision-making process, knowing it won't invent arbitrary functions or act outside the defined operational boundaries we set for it.

Tom: So these technical fixes are really impressive, but they aren’t all about fixing every single issue in one go; they’re building a more robust system piece by piece.

Jane: That makes sense; you can't fix everything at once, and focusing on the most structurally important components first is a smart strategy for building reliable AI.

Conclusion: Tom: To wrap up our discussion of "Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer," we’ve covered the key points about how this model was built and what it actually achieved in terms of performance across different scripts.

Jane: The overall message is that evaluating these models must be more sophisticated, especially when looking at low-resource languages where the standard benchmarks fail to capture true ability for Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Lu: I think the most exciting thing is how this demonstrates that we can achieve high performance by fixing structural issues like tokenization, rather than just trying to scale up indefinitely in the way of Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Meng: It proves that we can build a highly efficient and trustworthy model using clear methods like the tool-use contract and strong hardware utilization for real systems.

Lalam: It gives us hope for all the languages that aren't English by showing what is achievable when we are dedicated to building models like Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer to serve the community.

Tom: We’ve heard some brilliant insights today, but it’s time to wrap up our discussion of Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer.

Jane: Thank you all for joining us on the show and thank you for listening. We'll be back next week with another fascinating paper to discuss.

Qwen Team

cs.CL, cs.AI

Submitted: 2026-08-30

Updated: 2026-09-01

Comments: 7 pages, 6 tables. Companion paper (tokenizer): arXiv:2608.26449

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: The paper details the development and rigorous evaluation of Arkios, an open bilingual English-Nepali language model trained from scratch.

Key concepts

Devanagari-Aware Tokenizer
This is a custom tokenizer integrated into the model architecture to correctly handle Nepali script. It solves a structural problem where traditional methods break words at vowel signs, allowing the model to process complex scripts much better than standard methods.
Trained From Scratch
Arkios was built entirely from scratch rather than being fine-tuned on existing models. This involved massive upfront work in data curation and infrastructure setup, using 150B tokens of bilingual English-Nepali text to establish a solid baseline architecture.
Manifest-Conditioned Tool-Use Contract
This improvement restricts the model so it only calls a tool if the user explicitly declares that specific tool is available in the context. This prevents the model from learning to generate arbitrary functions, ensuring its behavior remains predictable and reliable for integration.

Terminology

Summary

The paper details the development and rigorous evaluation of Arkios, an open bilingual English-Nepali language model trained from scratch. This work is significant because it tackles the challenges inherent in low-resource language processing, providing a robust framework for Nepali generation that addresses common pitfalls such as script selection failures and domain-specific limitations.

Model Architecture and Training Stages

The evaluation compares two distinct checkpoints: the pretrained base model, which is prompted with few-shot completion (its native format), and the instruction-tuned chat model, which uses a zero-shot ChatML instruction format. The authors emphasize that each checkpoint is evaluated in the format it was actually trained to use. Furthermore, they note that No RLHF, DPO, or other preference optimization stage was applied; the released chat model is SFT-only. The model utilizes a 4096-token context window, which was fixed at pretraining time.

Evaluation of Target Script Fidelity

A primary focus of the evaluation is measuring the Fraction of translation outputs produced in the requested target script, as a fluent response in the wrong language can be misleadingly scored. The performance gap between models is stark:

  • The base model, prompted with naive few-shot completion, frequently continues generating in the source language rather than switching to the requested target. This is most severe when targeting Nepali, achieving only an 8% target-script rate en→ne.

  • The instruction-tuned checkpoint substantially improves this performance. The chat model achieves a 58% target-script rate for en→ne, while the base model reaches 81%. The authors conclude that this improvement indicates the base result is driven in large part by a prompting mismatch rather than solely by a property of the pretrained weights.

Technical Limitations and Reliability Constraints

The report identifies several critical limitations regarding reliability and capability. First, arithmetic and unit conversion are noted as unreliable without external assistance: No reliable arithmetic or unit conversion without a tool. This holds for the base model always, and for the chat model except where a matching tool is declared and actually invoked. Second, while the instruction-tuned checkpoint significantly closes the script gap, the remaining gap, particularly en→ne at 58%, remains a real limitation of the released chat model.

Data Constraints and Generalization Scope

The authors provide clear boundaries on the model's capabilities. Regarding Nepali proficiency, they state that Nepali capability is bounded by data availability, not architecture, noting that the entire practical public Nepali corpus available totaled approximately 4.18B unique tokens against a 342B-token English web pool. Furthermore, generalization claims must be handled with caution:

  • The ARC result is deemed likely domain-favorable... and should not be generalized into a broad capability claim.

  • The standard multiple-choice-letter evaluation format understates Nepali ability, requiring any third-party benchmarking to be read alongside this caveat.

Computational Methodology

To ensure the feasibility of these large evaluations, the methodology employed advanced memory management. The process involves gathering hidden states from the transformer body for an entire batch, applying the language-model head only at required token positions. This optimization is crucial because, given our vocabulary size (65,536) and typical batch–sequence-length products, it reduces the dominant memory allocation during scoring by approximately 40× relative to a naive full-logit implementation.

Improvements for AI systems

Based on this scientific paper excerpt, which details advanced evaluation methodologies for multilingual LLMs (specifically English Nepali), I have identified several critical limitations in current state-of-the-art AI systems. Implementing these improvements would significantly enhance reliability, especially in low-resource and multilingual settings.

Here are the specific architectural and methodological improvements required:

Improvement: Integrate a mandatory, dedicated Target Script/Language Gatekeeper Module immediately following the core generation transformer body. This module must operate as a post-processing constraint layer that explicitly penalizes or forces re-sampling of tokens belonging to the source language or any unintended script.

  • Mechanism: Instead of relying solely on prompt engineering (e.g., Generate in Nepali), the model must calculate the probability of generating a token t given that t belongs to the desired target script S target. If the probability mass for tokens outside S target exceeds a predefined threshold tau, a constrained beam search or rejection sampling mechanism must be triggered until compliance is achieved.

  • Improved Capability: The system can reliably produce outputs in the intended target script, even when the base model exhibits strong source-language continuation tendencies (e.g., achieving >95% target-script rate for en to ne), thereby eliminating languageselection failure from being misclassified as a translation quality failure.

  • Mechanism: The layer must analyze the metadata of the input format and adjust the weights or biases applied to the initial tokens, effectively simulating a format shift within its latent space. This ensures that zero-shot chat instructions are leveraged optimally, preventing degradation relative to base model performance when switching formats.

  • Improved Capability: The AI system can maintain peak performance regardless of whether it receives input via structured conversational prompts (ChatML) or through traditional few-shot completion examples, ensuring maximum consistency between the two operational modes and closing the performance gap observed in the current research.

  • Mechanism: The training objective must incorporate a loss function component that penalizes overly dominant reliance on high-resource languages when generating content for low-resource languages, forcing the model to learn deep structural representations specific to the limited data. Furthermore, the tokenizer must be optimized a priori for multilingual script boundaries rather than simply maximizing compression.

  • Improved Capability: The system can achieve significantly higher Nepali capability by efficiently utilizing smaller, specialized public corpora (e.g., 4.18B tokens) without suffering from catastrophic forgetting or being overwhelmed by the massive token count of high-resource English web pools (342B tokens).

  • Mechanism: When a tool is declared (e.g., for unit conversion), the system must internally verify that all units and parameters are compatible before generating any response text, and it must treat the tool's output as an immutable fact, overriding any contradictory textual generation attempts.

  • Improved Capability: The AI system can provide mathematically guaranteed outputs for arithmetic and unit conversions, eliminating the current unreliability observed when context requires external computation or specialized domain knowledge beyond general language patterns.

Abstract

We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single-file C/CUDA training stack and a Devanagari-aware byte-level BPE tokenizer built for this project. On ARC-Easy and ARC-Challenge, Arkios exceeds three comparably sized open models (Pythia-1.4B, TinyLlama-1.1B, OLMo-1B) despite an order of magnitude fewer training tokens, likely aided by a match between our educational-web-text pretraining data and ARC's grade-school-science format rather than a general capability advantage. We report full evaluation results under standard protocols, including a correction to an earlier partial-sample estimate, and findings specific to evaluating small models in a low-resource language: the standard multiple-choice-letter prompt format used by common evaluation harnesses places this model at chance on Nepali reading comprehension, and simultaneously at chance on English in the same format, which would lead a naive benchmark run to conclude the model has no Nepali ability when in fact it does. Concretely, both languages score at chance in the letter-choice format (0.240 Nepali, 0.236 English, against a chance baseline of 0.250), while scoring the answer text directly reveals genuine, English-favoring comprehension (0.306 Nepali, 0.387 English). We describe a manifest-conditioned tool-use contract introduced during instruction tuning, where tool calls are permitted only when a tool manifest is declared in context and suppressed otherwise, and report where that contract holds and where it does not. We release both the base and instruction-tuned model weights under Apache-2.0. The training code and a small privately-sourced portion of the Nepali pretraining corpus are not released; everything needed to reproduce the reported numbers from the released weights is included here.

Sources

Related papers