OPSD: an Offensive Persian Social media Dataset and its baseline evaluations

Mehran Safayani; Amir Sartipi; Amir Hossein Ahmadi; Parniyan Jalali; Amir Hossein Mansouri; Mohammad Bisheh-Niasar; Zahra Pourbahman

OPSD: an Offensive Persian Social media Dataset and its baseline evaluations

Mehran Safayani, Amir Sartipi, Amir Hossein Ahmadi, Parniyan Jalali, Amir Hossein Mansouri, Mohammad Bisheh-Niasar, Zahra Pourbahman

TL;DR

OPSD tackles the scarcity of Persian offensive language data by introducing a labeled OPSD dataset and a large unlabeled corpus for unsupervised learning. It employs a rigorous three-phase annotation protocol with five annotators and reports strong baseline performance from multilingual and Persian-specific transformer models, with additional gains from masked language modeling on unlabeled data. Key findings show XLM-RoBERTa as the top performer (NEG+ and 3-class) and demonstrate meaningful improvements from MLM, along with insightful error analysis that identifies label inconsistencies and data noise. The work lays a foundation for Persian offensive language detection with clear paths for scaling and preprocessing improvements to enhance real-world deployment.

Abstract

The proliferation of hate speech and offensive comments on social media has become increasingly prevalent due to user activities. Such comments can have detrimental effects on individuals' psychological well-being and social behavior. While numerous datasets in the English language exist in this domain, few equivalent resources are available for Persian language. To address this gap, this paper introduces two offensive datasets. The first dataset comprises annotations provided by domain experts, while the second consists of a large collection of unlabeled data obtained through web crawling for unsupervised learning purposes. To ensure the quality of the former dataset, a meticulous three-stage labeling process was conducted, and kappa measures were computed to assess inter-annotator agreement. Furthermore, experiments were performed on the dataset using state-of-the-art language models, both with and without employing masked language modeling techniques, as well as machine learning algorithms, in order to establish the baselines for the dataset using contemporary cutting-edge approaches. The obtained F1-scores for the three-class and two-class versions of the dataset were 76.9% and 89.9% for XLM-RoBERTa, respectively.

OPSD: an Offensive Persian Social media Dataset and its baseline evaluations

TL;DR

Abstract

OPSD: an Offensive Persian Social media Dataset and its baseline evaluations

Authors

TL;DR

Abstract

Table of Contents

Figures (5)