The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language
cs.CL
Submitted: 2026-08-25
Updated: 2026-08-25
Comments: 18 pages, 8 tables
Project page: https://w3techs.com/technologies/overview/content_language
License: http://creativecommons.org/licenses/by/4.0/
The gist: Persian (Farsi) is often described as a low-resource language in natural language processing, but that label collapses distinct shortages into a single category.
Terminology
Abstract
Persian (Farsi) is often described as a low-resource language in natural language processing, but that label collapses distinct shortages into a single category. This paper argues that Persian is more precisely described as annotation-scarce, provided that the term is understood as a property of its NLP resource ecology rather than an intrinsic property of the language. The review covers 34 representative Persian text resources available by July 2026 and adds three quantitative cross-checks. First, independent web measurements place Persian among roughly the twenty most visible content languages: W3Techs reports Persian on about 0.9% of websites with a known content language, while Common Crawl CC-MAIN-2026-30 identifies Persian as the primary language of 0.7039% of HTML pages. Second, a selective speech review shows a long resource trajectory from FARSDAT to recent corpora containing hundreds or thousands of hours of speech. Third, a matched Persian-English comparison normalizes task-specific annotation volumes by relative Common Crawl web presence. The resulting ratios vary sharply: Persian syntax and news NER are comparatively dense, whereas natural-language inference falls below the web-proportional baseline. The evidence therefore does not support a simple claim that Persian is globally deficient in labeled volume. Instead, annotation scarcity is expressed through uneven task and domain coverage, incompatible schemes, access and documentation friction, and limited supervision for specialist domains, preference data, and varieties beyond standard Iranian Persian.
Sources
- HmBlogs: A big general Persian corpus
- MIZAN: A Large Persian-English Parallel Corpus
- A Multi Purpose and Large Scale Speech Corpus in Persian and English for Speaker and Speech Recognition: the DeepMine Database
- ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
- PEYMA: A Tagged Corpus for Persian Named Entities
- SentiPers: A Sentiment Analysis Corpus for Persian
- ArmanEmo: A Persian Dataset for Text-based Emotion Detection
- Leveraging ParsBERT and Pretrained mT5 for Persian Abstractive Text Summarization
- Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
- FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models
- PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian
- PersianPunc: A Large-Scale Dataset and BERT-Based Approach for Persian Punctuation Restoration
- A hybrid entity-centric approach to Persian pronoun resolution
- Persian Rhetorical Structure Theory
- PersianMind: A Cross-Lingual Persian-English Large Language Model
- Benchmarking Open-Source Large Language Models for Persian in Zero-Shot and Few-Shot Learning
- Graphemic Normalization of the Perso-Arabic Script
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering