Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
cs.CL
Submitted: 2025-05-22
Updated: 2025-11-05
Comments: Accepted at AACL 2025. Camera-ready version
Journal ref: https://aclanthology.org/2025.ijcnlp-short.15/
DOI: 10.18653/v1/2025.ijcnlp-short.15
Code: https://github.com/grvkamath/low-resource-syn-ner
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Named Entity Recognition(NER) for low-resource languages aims to produce robust systems for languages where there is limited labeled training data available, and has been an area of increasing
Terminology
Abstract
Named Entity Recognition(NER) for low-resource languages aims to produce robust systems for languages where there is limited labeled training data available, and has been an area of increasing interest within NLP. Data augmentation for increasing the amount of low-resource labeled data is a common practice. In this paper, we explore the role of synthetic data in the context of multilingual, low-resource NER, considering 11 languages from diverse language families. Our results suggest that synthetic data does in fact hold promise for low-resource language NER, though we see significant variation between languages.
Sources
- GPT-4 Technical Report
- Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
- The Llama 3 Herd of Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering