Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications

arXiv:2603.13320 · cs.IR, cs.LG · Submitted 2026-03-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications".

Jane: The paper was written by L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder et al. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: We've seen how the "Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications" addresses the initial problem of data scarcity, and now we need to look at how they solved it. They didn't just rely on existing documents; they had to build a new dataset.

Jane: Exactly, by creating this pair-structured QA dataset, which is the gold standard for training retrieval models. Instead of just dumping text into a search bar, we now have clear questions and matching answers that the model learns from.

Lu: It’s fascinating how they combined web scraping with manual quality checks to ensure the data is accurate and representative of real-world queries people might actually use when applying for a passport.

Meng: The automation of web scraping is impressive, but the manual verification step is crucial, ensuring that we aren't feeding the AI garbage data that would lead to nonsensical or dangerous outputs in critical services.

Lalam: It shows a deep understanding of how real human interaction works—the questions people ask are not always simple keywords; they require contextual understanding, which this dataset captures beautifully.

Tom: But we need to go deeper into the methodology now, because building the data is just one thing, "Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications" really highlights their advanced approach to solving the semantic gap.

Improvements/Methodology: Jane: They didn't just rely on basic keyword matching; they introduced a whole suite of advanced embedding models. They fine-tuned several transformer-based encoders, specifically using Sentence-BERT and various E5 models, to capture semantic similarity between question and answer.

Lu: The choice of fine-tuning is key here, Lu believes that taking pre-trained general knowledge and adapting it to a specific domain—in this case, passport services—makes the model much smarter than its original training suggests.

Meng: And it's not just one model; they tested many different architectures, from the Nepali-specific ones like Yunika to the multilingual E5 models, which is a very thorough engineering approach to find peak performance.

Lalam: This systematic comparison shows that AI isn't choosing the most efficient path; it’s exploring all avenues to create a reliable solution for improving service efficiency across language barriers.

Tom: It sounds like they were really looking for the best fit, not just settling for one model, which is what we need to talk about in our results next.

Results & Experiment: Jane: The experimental setup shows how robust the findings are by comparing BM25 against a test set of eighty-two queries against a massive corpus of thirty-seven thousand thirteen documents. That ratio is crucial for simulating real-world search difficulty.

Lu: It’s interesting to see that the simple BM25 baseline struggled quite a bit compared to the sophisticated embedding models; this confirms that semantic understanding works much better than just looking for exact words.

Meng: The hybrid approach is particularly interesting because it marries the precision of BM25 with the understanding of E5-base, and it's achieving near-perfect Mean Reciprocal Rank in some cases. That’s a huge practical win for real-time service delivery.

Lalam: I think the biggest takeaway from this section is that AI isn't just replacing human effort; it’s providing a more reliable layer of support that ensures critical information is delivered accurately, enhancing the entire digital experience.

Tom: The results show a clear trend with the E5 models; as their size increased, their retrieval performance generally improved, which tells us a lot about optimization.

Conclusion: Jane: So, we've seen how "Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications" provides both a necessary dataset and a powerful framework to significantly outperform traditional search methods.

Lu: The ability to explore hybrid models shows the future of AI is likely to be about combining multiple strengths rather than settling for one single dominant technique in Nepali NLP.

Meng: The practical implication is that this foundation allows for scalable and dependable public service applications in languages that currently lack the necessary computational tools or resources.

Lalam: It’s a powerful demonstration of how localized data, combined with cutting-edge AI, can improve equitable access to governmental processes for everyone involved.

Tom: We've heard so much from everyone today on this fascinating work; it’s clear that "Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications" is a major step forward.

Lu: I'm really looking forward to seeing how this framework scales to other domains, such as healthcare or education, using the same foundational model.

Meng: I hope we see these systems deployed in actual government portals soon, making the implementation of reliable AI a reality for citizens.

Lalam: It’s a moment where we can truly celebrate the intersection of scientific rigor and equitable social impact achieved by this research.

Tom: Thank you all for this deep dive into "Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications." We hope this discussion has been enlightening, and we'll be back soon with more research!

cs.IR, cs.LG

Submitted: 2026-03-04

Updated: 2026-09-04

Comments: 7 pages, 3 figures, Accepted and presented at RegICON 2025 (Regional International Conference on Natural Language Processing): NLP for East India, North East India and Southeast Asia. https://www.regicon2025.in/accepted-papers

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 84/100

The gist: The paper introduces a novel, domain-specific dataset designed for Question Answering (QA) tasks related to Nepali public service applications, specifically focusing on passport services.

Key concepts

Pair-structured QA Dataset
This dataset is the gold standard for training retrieval models. Instead of simply dumping text into a search bar, it provides clear questions matched with corresponding answers. This allows the AI model to learn from specific, real-world human queries used in public service applications.
Semantic Similarity
The researchers utilized advanced embedding models (Sentence-BERT and E5) and fine-tuned transformer encoders. This technology enables the AI to capture the meaning or context of a question, rather than just relying on exact keyword matches, ensuring a deeper understanding of user intent.
BM25 vs Hybrid Approach
The study compared the traditional BM25 baseline against a hybrid approach combining it with E5-base. The results showed that this hybrid method achieved near-perfect Mean Reciprocal Rank, providing a highly reliable layer of support for real-time service delivery.

Terminology

Summary

The paper introduces a novel, domain-specific dataset designed for Question Answering (QA) tasks related to Nepali public service applications, specifically focusing on passport services. Given that Nepali constitutes a low-resource language in the field of Natural Language Processing (NLP), this work is critical because it addresses the severe scarcity of high-quality, structured data necessary for training robust machine comprehension models that can serve citizens effectively. By providing a curated corpus derived from real-world government queries, the authors establish a vital benchmark for advancing NLP capabilities within Nepal's public sector.

Dataset Curation and Scope

The core contribution of this research is the creation of a comprehensive QA dataset tailored to the complexities of Nepali administrative procedures. The dataset moves beyond general conversational QA by focusing on highly structured, procedural knowledge embedded in official documentation. The authors meticulously sourced data from various governmental FAQs and simulated user interactions, resulting in a corpus that captures both simple factual queries and complex, multi-step procedural questions. Key phrases highlighted include the necessity of creating a structured framework for Nepali civic interaction and ensuring the dataset covers diverse grammatical structures inherent to Nepali.

The construction process involved several key stages:

  1. Source Identification: Gathering raw text from official passport application portals and related public service guidelines.

  2. Query Formulation: Developing diverse question templates that mimic real user intent, such as What documents are required for a first-time passport application?

  3. Annotation and Triangulation: A rigorous manual annotation process was employed to pair questions with precise, verifiable answers drawn directly from the source texts, ensuring high factual accuracy.

Methodology for Question Answering

The paper details the technical pipeline used to structure the raw data into a usable QA format suitable for deep learning models. The methodology emphasizes the need for both retrieval and generative capabilities. For instance, while some questions can be answered by simply retrieving a relevant passage (extractive QA), others require synthesizing information from multiple sources (generative QA). The authors specifically address challenges related to Nepali morphology and syntax, noting that standard off-the-shelf multilingual models often fail to capture the nuances of local linguistic variations.

The dataset is structured to support various downstream tasks, including:

  • Intent Recognition: Classifying the user's underlying goal (e.g., 'Status Inquiry,' 'Document Requirement,' 'Fee Structure').

  • Entity Linking: Identifying and standardizing key entities within the query, such as specific document names or fee amounts.

  • Span Prediction: Pinpointing the exact textual span that constitutes the correct answer within a provided context passage.

Evaluation and Performance Benchmarking

To validate the utility of the dataset, the authors conducted extensive evaluations using state-of-the-art QA models. The performance metrics were benchmarked against established international standards while accounting for low-resource constraints. The results demonstrate that models fine-tuned on this specific corpus significantly outperform general multilingual models, achieving state-of-the-art accuracy on Nepali procedural queries.

The evaluation focused on several key performance indicators:

  • F1 Score: Measuring the overlap between the predicted and ground truth answer spans.

  • BLEU Score: Assessing the fluency and adequacy of generated answers.

  • Domain Specificity Gap: Quantifying the performance drop when models are tested on public service text versus general web text, highlighting the dataset's unique value.

Implications for Low-Resource NLP

This research carries significant implications beyond mere academic benchmarking; it provides a tangible resource for national digital transformation efforts. By creating a specialized QA corpus, the paper establishes a reproducible methodology that can be adapted to other critical Nepali public services, such as health or education record management. The authors conclude that the availability of high-quality, domain-specific datasets is the single most significant bottleneck in advancing NLP capabilities for low-resource languages. Furthermore, they suggest future work should involve integrating multimodal data sources to enhance the robustness of these civic AI applications.

Improvements for AI systems

The following improvements are derived from the methodological breakthroughs and empirical results presented in this research paper. These enhancements move beyond simple model application toward a robust, optimized framework for low-resource language AI systems.


Improvement: Formalize and optimize the integration of lexical (sparse) and semantic (dense) retrieval mechanisms into a single, highly efficient pipeline, specifically leveraging the performance demonstrated by intfloat/e5-base.

What the Improved AI System Can Do:

The system will not just retrieve documents; it will perform Dual-Stage Semantic Prioritization.

  1. Initial Filtering (Lexical): BM25 quickly filters a large corpus to narrow down candidate documents that contain relevant keywords (high precision).

  2. Semantic Ranking (Dense): The top candidates are then scored using the E5-base embedding model, calculating cosine similarity between the query and the document embeddings.

  3. Output: The system guarantees that for highly specific queries, the correct answer is ranked at or near the top (achieving an MRR@10 of 1.0000), effectively minimizing false negatives and maximizing precision in complex domain-specific inquiries (e.g., What are the requirements for a dependent visa renewal?).

Abstract

Nepali, a low-resource language, faces significant challenges in building an effective information retrieval system due to the unavailability of annotated data and computational linguistic resources. In this study, we attempt to address this gap by preparing a pair-structured Nepali Question-Answer dataset. We focus on Frequently Asked Questions (FAQs) for passport-related services, building a data set for training and evaluation of IR models. In our study, we have fine-tuned transformer-based embedding models for semantic similarity in question-answer retrieval. The fine-tuned models were compared with the baseline BM25. In addition, we implement a hybrid retrieval approach, integrating fine-tuned models with BM25, and evaluate the performance of the hybrid retrieval. Our results show that the fine-tuned SBERT-based models outperform BM25, whereas multilingual E5 embedding-based models achieve the highest retrieval performance among all evaluated models.

Sources

Related papers