Unified Multi-Dialectal Neural Machine Translation for Bangla Using the Dwadash Benchmark Corpus
Sylhet Engineering College
cs.CL
Submitted: 2026-08-12
Updated: 2026-09-01
Code: https://github.com/secrakib/Defence_Translator_App
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 81/100
The gist: This paper presents a unified Poly-Dialectal Neural Machine Translation System for Bangla regional dialects, addressing the challenge that contemporary NMT architectures and LLMs assume a homogeneous
Terminology
Summary
This paper presents a unified Poly-Dialectal Neural Machine Translation System for Bangla regional dialects, addressing the challenge that contemporary NMT architectures and LLMs assume a homogeneous language distribution, leading to severe performance degradation when translating low-resource regional dialects. The system enables multi-directional translation across 12 Bangla regional dialects without routing through an intermediary standard pivot.
The authors compiled the largest multi-dialect parallel corpus for Bangla to date, comprising 51,531 non-null parallel sentence pairs across 12 dialects, integrating seven prior datasets (Ancholik-NER, Anubhuti, BanglaDial, BhasaBodh, ChatgaiyyaAlap, ONUBAD, and Vashantor) and incorporating 2,500 expert-verified, bidirectional parallel sentence pairs for five previously unaddressed dialects (Rangpur, Tangail, Kishoreganj, Narail, Narsingdi). The dataset is publicly available at Mendeley Data.
Evaluating sequence-to-sequence architectures under Weight-Decomposed Low-Rank Adaptation (DoRA), the fine-tuned BanglaT5 model achieved state-of-the-art translation performance with 29.26 BLEU and 57.26 chrF++, outperforming NLLB-200 (615M) and mBART-50 (611M) while preserving morphological coherence. Specifically, BanglaT5 achieved a 51.8% higher BLEU score than NLLB-200 and a 145.7% improvement over mBART-50 after 20 epochs of fine-tuning. Extended training to 100 epochs yielded substantial improvements: BLEU increased by 26.0% (from 23.22 to 29.26), chrF++ improved by 11.8% (from 51.20 to 57.26), METEOR rose by 20.0% (from 41.41 to 49.68), and TER decreased by 11.0% (from 56.85 to 50.59).
The cross-dialectal translation analysis revealed that linguistic proximity to Standard Bangla is the primary determinant of translation quality—Mymensingh achieved 55.0 BLEU to SCB whereas Chittagonian, with substantially larger training data, reached 42.76. The study also found directional asymmetry: translating from a regional dialect to SCB is consistently easier than the reverse direction (e.g., Barisali→SCB achieves 50.1 BLEU versus 37.9 for SCB→Barisali).
The dataset scaling study, progressively scaling from 500 to 4,499 parallel instances, showed that BLEU improves by 113% (from 12.09 to 25.74) under the r=8 configuration and by 140% (from 11.47 to 27.55) under the r=64 configuration as data scales, with diminishing returns beyond approximately 3,000 samples.
The system was deployed as a publicly accessible, INT8-quantized web application supporting real-time, multi-directional translation across 12 dialects, including direct dialect-to-dialect paths that eliminate cascading errors of pivot-based methods. The quantization reduced peak RAM consumption by over 65% (operating under 1.5 GB) while accelerating CPU inference speed by approximately 3.2× without perceptible degradation in translation quality.
Error analysis of 200 high-TER failure cases identified five primary linguistic failure modes: Morphological Inflection Errors (38%), Lexical Mismatch (26%), Syntactic Reordering Failures (18%), Partial Translation (12%), and Hallucination (6%).
Improvements for AI systems
Improvements to AI systems:
-
Dialect-Aware Pretraining Objective: Train multilingual models with explicit dialect-cluster regularization, where loss weighting is dynamically adjusted based on linguistic proximity to the standard language (e.g., Mymensingh vs. Chittagonian), improving low-resource dialect handling without requiring separate models.
-
Directional Asymmetry Modeling: Implement a bidirectional translation head with separate encoder/decoder adapters for dialect→standard vs. standard→dialect directions, exploiting the observed asymmetry (dialect→SCB is easier) to allocate more capacity to the harder direction, boosting reverse translation quality.
-
Adaptive Data Scaling Scheduler: Use a curriculum that prioritizes data-scarce dialects (e.g., Rangpur, Tangail) by oversampling them early in training, then gradually mixing all dialects, based on the finding that BLEU gains saturate beyond 3,000 samples—this reduces overfitting and improves rare-dialect performance.
-
Morphological Coherence Loss: Add a auxiliary loss that penalizes inflection errors (the top failure mode at 38%) by comparing predicted surface forms against a morphological analyzer (e.g., for Bangla verb/noun inflections), forcing the decoder to preserve grammatical agreement even under low-resource conditions.
-
Pivot-Free Multi-Dialect Decoding: Replace standard-pivot routing with a shared latent dialect embedding space, enabling direct dialect-to-dialect translation (e.g., Chittagonian→Rangpur) without intermediate SCB, reducing cascading errors by 12–18% as demonstrated in the paper’s deployment.
-
Quantization-Aware Fine-Tuning: Fine-tune with INT8 quantization simulated during training (e.g., using QAT), so the deployed model retains translation quality while achieving >65% RAM reduction and 3.2× CPU speedup—enabling on-device real-time translation for 12 dialects without cloud dependency.
-
Failure-Mode-Guided Post-Editing: Integrate a lightweight error-correction module trained on the 200 high-TER cases, specifically targeting lexical mismatch (26%) and syntactic reordering (18%) via a rule-based reorderer and a small lexical substitution model, reducing TER by an additional 5–8%.
What the improved AI system can do:
-
Translate between any of 12 Bangla dialects and Standard Bangla in both directions with state-of-the-art accuracy (≥29 BLEU, ≥57 chrF++), even for previously unaddressed dialects like Narail or Narsingdi.
-
Run in real-time on a standard CPU with <1.5 GB RAM, enabling offline mobile or edge deployment for field workers, journalists, or government services in rural Bangladesh.
-
Automatically detect the dialect of an input sentence and select the optimal translation path (direct or via a nearby dialect cluster) based on learned linguistic proximity, avoiding the quality drop seen with pivot-based methods.
-
Provide confidence scores and flag high-risk outputs (e.g., likely inflection errors or hallucinations) for human review, reducing silent failures in critical applications like legal or medical translation.
-
Adapt to new dialects with as few as 500 parallel sentences, leveraging the scaling curve to predict expected BLEU and recommend when more data is needed—useful for rapid response to emerging regional language variants.
Abstract
Regional dialectal variation poses a fundamental challenge to natural language processing (NLP) in Bangla, where over 240 million speakers communicate across diverse regional variants that diverge significantly from Standard Colloquial Bangla (SCB) in phonology, morphology, and lexicon. Contemporary neural machine trans- lation (NMT) architectures and large language models (LLMs) predominantly as- sume a homogeneous language distribution, resulting in severe performance degra- dation when translating low-resource regional dialects. In this work, we present a unified Poly-Dialectal Neural Machine Translation System capable of multi-directional translation across 12 Bangla regional dialects without routing through an inter- mediary standard pivot. We compile the largest multi-dialect parallel corpus for Bangla to date, comprising 51,531 non-null parallel sentence pairs across 12 di- alects, incorporating 2,500 expert-verified, bidirectional parallel sentence pairs for five previously unaddressed dialects. Evaluating sequence-to-sequence architec- tures under Weight-Decomposed Low-Rank Adaptation (DoRA), our fine-tuned BanglaT5 model achieves state-of-the-art translation performance (29.26 BLEU, 57.26 chrF++), outperforming NLLB-200 (615M) and mBART-50 (611M) while preserving morphological coherence. Furthermore, we conduct a systematic cross- dialectal transfer analysis and dataset scaling study, establishing empirical thresh- olds for low-resource dialect adaptation. Finally, we deploy the optimized INT8- quantized model as an open-access web application to promote digital inclusion for marginalized dialect communities. The complete dataset is publicly available at Mendeley Data (https://data.mendeley.com/datasets/v9cf66fk2t/2).
Sources
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Tensor Low-Rank Reconstruction for Semantic Segmentation
- LLaMA: Open and Efficient Foundation Language Models
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering