ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

arXiv:2609.16393 · cs.CL · Submitted 2026-09-14 · Read on arXiv

cs.CL

Submitted: 2026-09-14

Updated: 2026-09-14

Comments: Accepted to EMNLP 2026 (Main Conference)

Code: https://github.com/zbokaee/ParsHate

License: http://creativecommons.org/licenses/by/4.0/

The gist: We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian.

Terminology

Abstract

We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and supports both hate detection and multi-label fine-grained target identification across seven structured target categories. ParsHate also distinguishes explicit and implicit hate, marks explicit and implicit targets, and provides span-level rationales. Data collection combines random and score-stratified temporal sampling to reduce keyword-driven bias while preserving natural label distributions. Applying SOTA models for Persian hate-speech detection on ParsHate shows moderate performance (79% F1), especially with samples from earlier years, and low performance with target identification (25.5% macro-F1). This emphasizes the diverse sampling of hate speech in ParsHate and its challenging nature that requires more advanced methods for better performance. Dataset is made publicly available.

Sources

Related papers