Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam

arXiv:2608.29959 · cs.CL · Submitted 2026-08-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Generative vs. Encoder Models for Multilingual NER".

Tom: This paper presents a "rigorous comparative study of generative and encoder-based neural architectures for NER" across the eleven languages of the Naamapadam benchmark.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: We’ve seen how generative models often fall short when trying to handle complex, structured data, so let’s look at the practical recommendations offered by this study regarding "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam."

Jane: The authors didn't just stop at identifying where the problems were; they provided a clear and actionable roadmap for fixing them, which is incredibly helpful when we are designing complex multilingual systems.

Lu: I find their specific advice regarding cross-lingual transfer and adaptation techniques particularly exciting because it suggests a way to adapt general knowledge to solve local data scarcity issues instead of starting from scratch.

Meng: For implementation, the recommendation of using an entity-aware hybrid sampling strategy is a massive practical win—it gives us a direct, mechanical way to balance our training data for specialized organizational entities that are usually ignored.

Lalam: It’s extremely encouraging that they provide this level of technical detail; it means we have a clear roadmap to build systems that are designed not just to work, but to be fundamentally equitable across all linguistic communities.

Tom: The authors guide us into operational planning, suggesting specific model choices based on whether the language falls into the Encoder-Dominant or Partial Coverage cluster.

Jane: It’s crucial for our listeners to understand that these clusters aren't just academic labels; they are real-world categories that require different engineering approaches to meet performance goals.

Lu: And Meng is right, we need to consider how complex this advice is; using adapters and transfer learning sounds like a massive undertaking for implementation in production environments.

Meng: It is a significant effort, but the guidance on the Failure Zone—like insisting on expanding script-level tokenizers before any modeling—gives us a clear priority list that makes sense when dealing with extremely sparse data.

Lalam: This careful, tailored approach ensures that when AI struggles with one specific community's data, we are providing them with a practical way to achieve digital inclusion through targeted engineering.

Tom: The recommendations for the "Failure Zone" language of Oriya are particularly strong, given the near-zero performance and extremely sparse entity vocabulary.

Jane: It’s clear that by offering these targeted strategies, the authors have provided a blueprint for building reliable, culturally sensitive AI systems instead of just hoping for some universal fix.

Lu: This approach acknowledges the limits of scaling and suggests that true innovation comes from focusing on localized data gaps rather than relying on general knowledge alone.

Meng: I’m particularly interested in how this strategy addresses the class imbalance issue; it directly tackles the fact that non-entity tokens constitute such a high percentage of training data.

Lalam: This methodical approach ensures that AI can help bridge the gap between technological advancement and societal equity for communities that have been underserved.

The paper's summary: Tom: We’ve seen how generative models often struggle with structured tasks, so let's look at the final conclusion of our discussion on "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam."

Jane: It's clear that building truly equitable AI requires a highly nuanced approach rather than a one-size-fits-all solution that ignores linguistic complexity.

Lu: It’s a powerful reminder that technical prowess must always be guided by deep knowledge of the communities we aim to serve digitally, understanding their specific needs.

Meng: And for developers, this means our focus needs to shift from chasing general benchmarks toward engineering robust, targeted solutions for specific language clusters based on the evidence.

Lalam: Ultimately, this research gives us a tangible blueprint—a roadmap—for achieving digital inclusion across the most linguistically diverse regions of the world.

Tom: I think what really stands out is the practical takeaway: we now have proven architectural paths forward for every scenario, from Encoder-Dominant to Failure Zone languages.

Jane: It’s a monumental step forward because it moves us from identifying problems to prescribing precise, actionable engineering remedies for each one.

Lu: I hope this work inspires more creative solutions for tackling the inherent biases in language processing across all future research endeavors globally.

Meng: My hope is that these targeted recommendations are finally used by developers who need truly reliable AI systems for production deployment, not just academic models.

Lalam: Lalam hopes this leads to a world where digital infrastructure serves every person, regardless of their linguistic background or the scarcity of their local data.

The paper's improvements: Tom: So, wrapping up our deep dive into "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam," it's clear that building truly equitable AI requires a highly nuanced, multilingual approach rather than a one-size-fits-all solution.

Jane: Exactly; the study has taught us that recognizing linguistic structure and addressing local data scarcity are far more critical to success than simply increasing the size of the underlying models.

Lu: It’s a powerful reminder that technical prowess must always be guided by deep anthropological understanding of the communities we aim to serve digitally.

Meng: And for developers, this means our focus needs to shift from chasing general model performance benchmarks toward engineering robust, targeted solutions for specific language clusters based on the evidence.

Lalam: Ultimately, this research gives us a tangible blueprint—a roadmap—for achieving digital inclusion across the most linguistically diverse regions of the world.

Tom: I think what really stands out is the practical takeaway: we now have proven architectural paths forward for every scenario, from Encoder-Dominant to Failure Zone languages.

Jane: It’s a monumental step forward because it moves us from identifying problems to prescribing precise, actionable engineering remedies for each one.

Lu: I hope this work inspires more creative solutions for tackling the inherent biases in language processing across all future research endeavors.

Meng: My hope is that these targeted recommendations are finally used by developers who need truly reliable AI systems for production deployment, not just academic models.

Lalam: Lalam hopes this leads to a world where digital infrastructure serves every person, regardless of their linguistic background or the scarcity of their local data.

Tom: Thank you all for joining us today; we've had an incredibly informative look at the complexities unveiled by "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam."

Jane: We have a lot to digest from this one, but we're ready to pivot our focus next and examine how these architectural lessons might apply to the challenges of low-resource image recognition.

Conclusion: Tom: So, to wrap up our deep dive into "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam," it’s really clear that building equitable AI demands a highly nuanced approach, far from any simple one-size-fits-all solution.

Jane: Exactly; the biggest takeaway is that understanding the linguistic structure and addressing local data scarcity are exponentially more critical to success than simply chasing bigger, more general models.

Lu: It’s a powerful reminder that technical prowess must always be guided by deep anthropological understanding of the communities we aim to serve digitally.

Meng: And for developers listening, this means our focus has to pivot away from chasing general benchmarks toward engineering robust, targeted solutions for specific language clusters based on the evidence presented.

Lalam: Ultimately, this research provides us with a tangible blueprint—a true roadmap—for achieving digital inclusion across the most linguistically diverse regions of the world.

Tom: I think what truly stands out is the practical takeaway: we now have proven architectural paths forward for every scenario, from Encoder-Dominant to those challenging Failure Zone languages.

Jane: It really is a monumental step forward because it moves us from just identifying problems to prescribing precise, actionable engineering remedies for each one of them.

Lu: I hope this work inspires more creative solutions for tackling the inherent biases in language processing across all future research endeavors globally.

Meng: My hope is that these targeted recommendations are finally adopted by developers who need truly reliable AI systems for real-world production deployment, not just academic models.

Lalam: Lalam genuinely hopes this leads to a world where digital infrastructure serves every single person, regardless of their linguistic background or the scarcity of their local data.

Tom: Thank you all for joining us today; it’s been an incredibly informative look at the complexities unveiled by "Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam."

Jane: We have a lot to digest from this one, but we are ready to pivot our focus next and examine how these architectural lessons might apply to the challenges of low-resource image recognition.

Jakkala Mahesh, Jatavath Shravan Kumar, Komalla Shivani, Sujoy Sarkar

Rajiv Gandhi University of Knowledge Technologies, Basar, Telangana, India

cs.CL

Submitted: 2026-08-30

Updated: 2026-08-30

Code: https://github.com/MaheshJakkala/naamapadam-multilingual-ner

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: This paper presents a "rigorous comparative study of generative and encoder-based neural architectures for NER" across the eleven languages of the Naamapadam benchmark.

Key concepts

Generative vs. Encoder Models
The study compares generative models against encoder-based neural architectures when performing Named Entity Recognition (NER). The discussion focuses on where generative models fall short with structured data and how encoder models perform in this context.
Naamapadam Benchmark
This is the specific dataset used in the empirical study. It consists of eleven languages, which serves as the testbed for comparing generative and encoder models in a multilingual NER task.
Encoder-Dominant vs. Partial Coverage Cluster
These are real-world categories derived from the study that require different engineering approaches. The choice between these clusters dictates the specific model architecture and strategy needed to meet performance goals for a given language.
Failure Zone
This refers to languages, like Oriya, that showed near-zero performance and extremely sparse entity vocabulary. The study recommends specific actions for these zones, such as expanding script-level tokenizers before modeling.

Terminology

Summary

This paper presents a rigorous comparative study of generative and encoder-based neural architectures for NER across the eleven languages of the Naamapadam benchmark. It addresses the resource poverty and digital layer incompleteness of Indic languages, investigating whether emerging Large Language Models (LLMs) can replace painstaking annotation in low-resource Named Entity Recognition tasks.

Experimental Methodology

The researchers conducted a comprehensive evaluation under strict CoNLL span-level evaluation protocols, comparing five classic model families, four decoder-only LLMs fine-tuned via LoRA and 4-bit NF4 quantization, and nine generative models using zero-to-5-shot inference. To combat chronic ORG class-imbalance, they implemented an entity-aware hybrid sampling strategy consisting of a 50/50 split between uniform random sampling and stratified entity enrichment.

The experimental setup included the following:

  • Five classic families: T5, FLAN-T5, mT5, mBERT, and XLM-R.

  • Four decoder-only LLMs: TinyLlama, LLaMA-3.2, Gemma-2, and Qwen2.5.

  • Nine generative models in zero-to-5-shot inference on Hindi.

Due to GPU constraints, fine-tuning was limited to 5,000 training and 500 testing samples per language.

Key Empirical Findings

The study reveals a parameter-efficiency inversion, where encoder-based models with only 177–278M parameters substantially outperform every generative architecture in ten of eleven languages. The performance gaps range from 7.5 to 40 percentage points against the strongest competitor, Gemma-2-2B, and the best few-shot result reaches only 28% of the encoder baseline.

The researchers identified three distinct language clusters:

  1. Encoder-Dominant: Languages like Hindi, Bengali, and Tamil where encoders achieve F1 > 0.50 and trail generative models by 14–40 pp.

  2. Partial-Coverage: Assamese, where Gemma-2 shows approximate parity with encoders, suggesting broad multilingual pre-training partially compensates for sparse data.

  3. Failure-Zone: Oriya, where all generative models achieved near-zero F1 due to Odia script tokenisation gaps and extremely sparse entity vocabulary.

Error Analysis and Structural Bottlenecks

The researchers concluded that architectural suitability; not parameter count or pre-training scale, governs strict-evaluation NER performance. A primary bottleneck for generative models is format compliance, specifically BIO malformation, which accounts for 55–80% of errors in seq2seq models and 18–34% in decoder LLMs. Unlike generative models, encoder-based classification heads are architecturally immune to these structural errors.

Furthermore, decoder-only LLMs exhibit a 22% boundary off-by-one errors rate compared to 12% for encoders, arising from left-to-right generation bias. Each such error carries a 2× F1 penalty under exact-match evaluation, further widening the performance gap between architectures.

Deployment Recommendations

The paper provides actionable deployment guidelines grounded in transfer learning and low-resource NLP principles. For production-quality Indic NLP pipelines, the authors recommend the following strategies:

  • Deploy mBERT or XLM-R with entity-aware hybrid sampling for encoder-dominant languages.

  • Utilize cross-lingual transfer from Bengali plus entity substitution augmentation for partial-coverage languages like Assamese.

  • Prioritize script-level tokeniser extension and targeted data collection for constrained-resource zones like Oriya, as these must precede modelling efforts.

Improvements for AI systems

1. Implementation of Constrained Autoregressive Decoding

  • Improvement: Integrate a masking layer during the inference phase of decoder-only LLMs to enforce valid BIO (Beginning, Inside, Outside) sequence transitions and prevent illegal label shifts (e.g., preventing an I-ORG tag from following a B-LOC tag).

  • Capability: The AI system will eliminate BIO malformation errors and significantly reduce boundary off-by-one errors, ensuring 100% structural compliance with strict CoNLL span-level evaluation standards.

2. Multi-Tiered Language-Specific Model Routing

  • Improvement: Deploy a language-detection router that directs queries to specific architectures based on identified linguistic clusters: (a) high-performance, low-parameter Encoders (XLM-R/mBERT) for Encoder-Dominant languages; (b) instruction-tuned LLMs with cross-lingual transfer (e.g., Bengali-to-Assamese) for Partial-Coverage languages; and (c) specialized script-extended models for Failure-Zone languages.

  • Capability: The AI system will maintain high NER accuracy across diverse linguistic landscapes, preventing the catastrophic performance drops (up to 40+ percentage points) currently seen when using generic LLMs for morphologically complex or script-unique languages.

3. Entity-Aware Stratified Fine-Tuning Pipeline

  • Improvement: Replace uniform random sampling in the fine-tuning stage with a 50/50 hybrid sampling strategy: 50% uniform random sampling and 50% stratified entity enrichment that guarantees a minimum threshold of examples (e.g., 50) for minority classes like ORG.

  • Capability: The AI system will achieve significantly higher recall for under-represented entity types, overcoming the class imbalance caused by the overwhelming dominance of "O" (non-entity) tokens in news-domain corpora.

4. Script-Specific Vocabulary Extension for Generative Models

  • Improvement: For Failure-Zone languages (e.g., Oriya), implement a pre-fine-tuning stage involving vocabulary expansion and subword-level training on the generative model's tokenizer to bridge script-coverage gaps.

  • Capability: The AI system will cease tokenization-level hallucinations—where the model outputs Roman or Hindi characters in place of target-language glyphs—enabling reliable generative NER in low-resource, script-unique environments.

5. Instruction-Tuned Few-Shot Optimization

  • Improvement: Prioritize the use of instruction-tuned models (e.g., Gemma-IT) over base models for few-shot NER tasks, specifically optimizing the prompt structure to reinforce the token(LABEL) format.

  • Capability: The AI system will demonstrate a 3.8× improvement in few-shot NER performance, allowing for rapid deployment in new languages where only a handful of annotated examples are available.

Sources

Related papers