Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Generative vs. Encoder Models for Multilingual NER".
Tom: This paper presents a "rigorous comparative study of generative and encoder-based neural architectures for NER" across the eleven languages of the Naamapadam benchmark.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: We’ve seen how generative models often fall short when trying to handle complex, structured data, so let’s look at the practical recommendations offered by this study regarding "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam."
Jane: The authors didn't just stop at identifying where the problems were; they provided a clear and actionable roadmap for fixing them, which is incredibly helpful when we are designing complex multilingual systems.
Lu: I find their specific advice regarding cross-lingual transfer and adaptation techniques particularly exciting because it suggests a way to adapt general knowledge to solve local data scarcity issues instead of starting from scratch.
Meng: For implementation, the recommendation of using an entity-aware hybrid sampling strategy is a massive practical win—it gives us a direct, mechanical way to balance our training data for specialized organizational entities that are usually ignored.
Lalam: It’s extremely encouraging that they provide this level of technical detail; it means we have a clear roadmap to build systems that are designed not just to work, but to be fundamentally equitable across all linguistic communities.
Tom: The authors guide us into operational planning, suggesting specific model choices based on whether the language falls into the Encoder-Dominant or Partial Coverage cluster.
Jane: It’s crucial for our listeners to understand that these clusters aren't just academic labels; they are real-world categories that require different engineering approaches to meet performance goals.
Lu: And Meng is right, we need to consider how complex this advice is; using adapters and transfer learning sounds like a massive undertaking for implementation in production environments.
Meng: It is a significant effort, but the guidance on the Failure Zone—like insisting on expanding script-level tokenizers before any modeling—gives us a clear priority list that makes sense when dealing with extremely sparse data.
Lalam: This careful, tailored approach ensures that when AI struggles with one specific community's data, we are providing them with a practical way to achieve digital inclusion through targeted engineering.
Tom: The recommendations for the "Failure Zone" language of Oriya are particularly strong, given the near-zero performance and extremely sparse entity vocabulary.
Jane: It’s clear that by offering these targeted strategies, the authors have provided a blueprint for building reliable, culturally sensitive AI systems instead of just hoping for some universal fix.
Lu: This approach acknowledges the limits of scaling and suggests that true innovation comes from focusing on localized data gaps rather than relying on general knowledge alone.
Meng: I’m particularly interested in how this strategy addresses the class imbalance issue; it directly tackles the fact that non-entity tokens constitute such a high percentage of training data.
Lalam: This methodical approach ensures that AI can help bridge the gap between technological advancement and societal equity for communities that have been underserved.
The paper's summary: Tom: We’ve seen how generative models often struggle with structured tasks, so let's look at the final conclusion of our discussion on "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam."
Jane: It's clear that building truly equitable AI requires a highly nuanced approach rather than a one-size-fits-all solution that ignores linguistic complexity.
Lu: It’s a powerful reminder that technical prowess must always be guided by deep knowledge of the communities we aim to serve digitally, understanding their specific needs.
Meng: And for developers, this means our focus needs to shift from chasing general benchmarks toward engineering robust, targeted solutions for specific language clusters based on the evidence.
Lalam: Ultimately, this research gives us a tangible blueprint—a roadmap—for achieving digital inclusion across the most linguistically diverse regions of the world.
Tom: I think what really stands out is the practical takeaway: we now have proven architectural paths forward for every scenario, from Encoder-Dominant to Failure Zone languages.
Jane: It’s a monumental step forward because it moves us from identifying problems to prescribing precise, actionable engineering remedies for each one.
Lu: I hope this work inspires more creative solutions for tackling the inherent biases in language processing across all future research endeavors globally.
Meng: My hope is that these targeted recommendations are finally used by developers who need truly reliable AI systems for production deployment, not just academic models.
Lalam: Lalam hopes this leads to a world where digital infrastructure serves every person, regardless of their linguistic background or the scarcity of their local data.
The paper's improvements: Tom: So, wrapping up our deep dive into "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam," it's clear that building truly equitable AI requires a highly nuanced, multilingual approach rather than a one-size-fits-all solution.
Jane: Exactly; the study has taught us that recognizing linguistic structure and addressing local data scarcity are far more critical to success than simply increasing the size of the underlying models.
Lu: It’s a powerful reminder that technical prowess must always be guided by deep anthropological understanding of the communities we aim to serve digitally.
Meng: And for developers, this means our focus needs to shift from chasing general model performance benchmarks toward engineering robust, targeted solutions for specific language clusters based on the evidence.
Lalam: Ultimately, this research gives us a tangible blueprint—a roadmap—for achieving digital inclusion across the most linguistically diverse regions of the world.
Tom: I think what really stands out is the practical takeaway: we now have proven architectural paths forward for every scenario, from Encoder-Dominant to Failure Zone languages.
Jane: It’s a monumental step forward because it moves us from identifying problems to prescribing precise, actionable engineering remedies for each one.
Lu: I hope this work inspires more creative solutions for tackling the inherent biases in language processing across all future research endeavors.
Meng: My hope is that these targeted recommendations are finally used by developers who need truly reliable AI systems for production deployment, not just academic models.
Lalam: Lalam hopes this leads to a world where digital infrastructure serves every person, regardless of their linguistic background or the scarcity of their local data.
Tom: Thank you all for joining us today; we've had an incredibly informative look at the complexities unveiled by "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam."
Jane: We have a lot to digest from this one, but we're ready to pivot our focus next and examine how these architectural lessons might apply to the challenges of low-resource image recognition.
Conclusion: Tom: So, to wrap up our deep dive into "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam," it’s really clear that building equitable AI demands a highly nuanced approach, far from any simple one-size-fits-all solution.
Jane: Exactly; the biggest takeaway is that understanding the linguistic structure and addressing local data scarcity are exponentially more critical to success than simply chasing bigger, more general models.
Lu: It’s a powerful reminder that technical prowess must always be guided by deep anthropological understanding of the communities we aim to serve digitally.
Meng: And for developers listening, this means our focus has to pivot away from chasing general benchmarks toward engineering robust, targeted solutions for specific language clusters based on the evidence presented.
Lalam: Ultimately, this research provides us with a tangible blueprint—a true roadmap—for achieving digital inclusion across the most linguistically diverse regions of the world.
Tom: I think what truly stands out is the practical takeaway: we now have proven architectural paths forward for every scenario, from Encoder-Dominant to those challenging Failure Zone languages.
Jane: It really is a monumental step forward because it moves us from just identifying problems to prescribing precise, actionable engineering remedies for each one of them.
Lu: I hope this work inspires more creative solutions for tackling the inherent biases in language processing across all future research endeavors globally.
Meng: My hope is that these targeted recommendations are finally adopted by developers who need truly reliable AI systems for real-world production deployment, not just academic models.
Lalam: Lalam genuinely hopes this leads to a world where digital infrastructure serves every single person, regardless of their linguistic background or the scarcity of their local data.
Tom: Thank you all for joining us today; it’s been an incredibly informative look at the complexities unveiled by "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam."
Jane: We have a lot to digest from this one, but we are ready to pivot our focus next and examine how these architectural lessons might apply to the challenges of low-resource image recognition.
Jakkala Mahesh, Jatavath Shravan Kumar, Komalla Shivani, Sujoy Sarkar
Rajiv Gandhi University of Knowledge Technologies, Basar, Telangana, India
cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Code: https://github.com/MaheshJakkala/naamapadam-multilingual-ner
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 84/100
The gist: This paper presents a "rigorous comparative study of generative and encoder-based neural architectures for NER" across the eleven languages of the Naamapadam benchmark.
Key concepts
- Generative vs. Encoder Models
- The study compares generative models against encoder-based neural architectures when performing Named Entity Recognition (NER). The discussion focuses on where generative models fall short with structured data and how encoder models perform in this context.
- Naamapadam Benchmark
- This is the specific dataset used in the empirical study. It consists of eleven languages, which serves as the testbed for comparing generative and encoder models in a multilingual NER task.
- Encoder-Dominant vs. Partial Coverage Cluster
- These are real-world categories derived from the study that require different engineering approaches. The choice between these clusters dictates the specific model architecture and strategy needed to meet performance goals for a given language.
- Failure Zone
- This refers to languages, like Oriya, that showed near-zero performance and extremely sparse entity vocabulary. The study recommends specific actions for these zones, such as expanding script-level tokenizers before modeling.
Terminology
Summary
This paper presents a rigorous comparative study of generative and encoder-based neural architectures for NER
across the eleven languages of the Naamapadam benchmark. It addresses the resource poverty
and digital layer
incompleteness of Indic languages, investigating whether emerging Large Language Models (LLMs) can replace painstaking annotation
in low-resource Named Entity Recognition tasks.
Experimental Methodology
The researchers conducted a comprehensive evaluation under strict CoNLL span-level evaluation
protocols, comparing five classic model families, four decoder-only LLMs fine-tuned via LoRA and 4-bit NF4 quantization, and nine generative models using zero-to-5-shot inference. To combat chronic ORG class-imbalance,
they implemented an entity-aware hybrid sampling strategy
consisting of a 50/50 split between uniform random sampling and stratified entity enrichment.
The experimental setup included the following:
-
Five classic families: T5, FLAN-T5, mT5, mBERT, and XLM-R.
-
Four decoder-only LLMs: TinyLlama, LLaMA-3.2, Gemma-2, and Qwen2.5.
-
Nine generative models in zero-to-5-shot inference on Hindi.
Due to GPU constraints, fine-tuning was limited to 5,000 training and 500 testing samples per language.
Key Empirical Findings
The study reveals a parameter-efficiency inversion,
where encoder-based models with only 177–278M parameters substantially outperform every generative architecture in ten of eleven languages.
The performance gaps range from 7.5 to 40 percentage points against the strongest competitor, Gemma-2-2B, and the best few-shot result reaches only 28% of the encoder baseline.
The researchers identified three distinct language clusters:
-
Encoder-Dominant: Languages like Hindi, Bengali, and Tamil where encoders achieve F1 > 0.50 and trail generative models by 14–40 pp.
-
Partial-Coverage: Assamese, where Gemma-2 shows
approximate parity
with encoders, suggestingbroad multilingual pre-training partially compensates for sparse data.
-
Failure-Zone: Oriya, where all generative models achieved
near-zero F1
due toOdia script tokenisation gaps
andextremely sparse entity vocabulary.
Error Analysis and Structural Bottlenecks
The researchers concluded that architectural suitability; not parameter count or pre-training scale, governs strict-evaluation NER performance.
A primary bottleneck for generative models is format compliance,
specifically BIO malformation,
which accounts for 55–80% of errors in seq2seq models and 18–34% in decoder LLMs. Unlike generative models, encoder-based classification heads are architecturally immune
to these structural errors.
Furthermore, decoder-only LLMs exhibit a 22% boundary off-by-one errors
rate compared to 12% for encoders, arising from left-to-right generation bias.
Each such error carries a 2× F1 penalty under exact-match evaluation,
further widening the performance gap between architectures.
Deployment Recommendations
The paper provides actionable deployment guidelines grounded in transfer learning and low-resource NLP principles.
For production-quality Indic NLP pipelines, the authors recommend the following strategies:
-
Deploy mBERT or XLM-R with
entity-aware hybrid sampling
for encoder-dominant languages. -
Utilize cross-lingual transfer from Bengali plus
entity substitution augmentation
for partial-coverage languages like Assamese. -
Prioritize
script-level tokeniser extension and targeted data collection
for constrained-resource zones like Oriya, as these mustprecede modelling efforts.
Improvements for AI systems
1. Implementation of Constrained Autoregressive Decoding
-
Improvement: Integrate a masking layer during the inference phase of decoder-only LLMs to enforce valid BIO (Beginning, Inside, Outside) sequence transitions and prevent illegal label shifts (e.g., preventing an
I-ORGtag from following aB-LOCtag). -
Capability: The AI system will eliminate
BIO malformation
errors and significantly reduceboundary off-by-one
errors, ensuring 100% structural compliance with strict CoNLL span-level evaluation standards.
2. Multi-Tiered Language-Specific Model Routing
-
Improvement: Deploy a language-detection router that directs queries to specific architectures based on identified linguistic clusters: (a) high-performance, low-parameter Encoders (XLM-R/mBERT) for
Encoder-Dominant
languages; (b) instruction-tuned LLMs with cross-lingual transfer (e.g., Bengali-to-Assamese) forPartial-Coverage
languages; and (c) specialized script-extended models forFailure-Zone
languages. -
Capability: The AI system will maintain high NER accuracy across diverse linguistic landscapes, preventing the catastrophic performance drops (up to 40+ percentage points) currently seen when using generic LLMs for morphologically complex or script-unique languages.
3. Entity-Aware Stratified Fine-Tuning Pipeline
-
Improvement: Replace uniform random sampling in the fine-tuning stage with a 50/50 hybrid sampling strategy: 50% uniform random sampling and 50% stratified entity enrichment that guarantees a minimum threshold of examples (e.g., 50) for minority classes like
ORG. -
Capability: The AI system will achieve significantly higher recall for under-represented entity types, overcoming the class imbalance caused by the overwhelming dominance of "O" (non-entity) tokens in news-domain corpora.
4. Script-Specific Vocabulary Extension for Generative Models
-
Improvement: For
Failure-Zone
languages (e.g., Oriya), implement a pre-fine-tuning stage involving vocabulary expansion and subword-level training on the generative model's tokenizer to bridge script-coverage gaps. -
Capability: The AI system will cease
tokenization-level hallucinations
—where the model outputs Roman or Hindi characters in place of target-language glyphs—enabling reliable generative NER in low-resource, script-unique environments.
5. Instruction-Tuned Few-Shot Optimization
-
Improvement: Prioritize the use of instruction-tuned models (e.g., Gemma-IT) over base models for few-shot NER tasks, specifically optimizing the prompt structure to reinforce the
token(LABEL)format. -
Capability: The AI system will demonstrate a 3.8× improvement in few-shot NER performance, allowing for rapid deployment in new languages where only a handful of annotated examples are available.
Sources
- Scaling Instruction-Finetuned Language Models
- The Llama 3 Herd of Models
- Gemma: Open Models Based on Gemini Research and Technology
- TinyLlama: An Open-Source Small Language Model
- Are Emojis Emotional? A Study to Understand the Association between Emojis and Emotions
- MuRIL: Multilingual Representations for Indian Languages
- DSMNet: Deep High-precision 3D Surface Modeling from Sparse Point Cloud Frames
- GPT-NER: Named Entity Recognition via Large Language Models
- Large Language Models: A Survey
- Bidirectional LSTM-CRF Models for Sequence Tagging
- A Survey of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering