Table of Contents
Fetching ...

Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media

Dennis Assenmacher, Paloma Piot, Katarina Laken, David Jurgens, Claudia Wagner

TL;DR

This work tackles the gap in dehumanization detection by introducing a large bilingual (English–German) dataset drawn from $\mathbb{X}$ and Reddit, containing over 490,000 candidate items and 16,000 annotated instances across animalistic/mechanistic and explicit/implicit dimensions. Using theory-informed sampling, expert crowd annotation, and span-level labeling, the authors demonstrate substantial performance gains for 14 modeling approaches, with notable improvements in zero- and few-shot settings when fine-tuned on the new data. The study also shows robust cross-lingual generalization and provides a detailed error analysis that informs future model improvements. By publishing the full candidate pool and associated predictions, the work offers a valuable resource for training, evaluation, and ongoing data collection in dehumanization detection with practical implications for online safety and bias mitigation.

Abstract

Digital dehumanization, although a critical issue, remains largely overlooked within the field of computational linguistics and Natural Language Processing. The prevailing approach in current research concentrating primarily on a single aspect of dehumanization that identifies overtly negative statements as its core marker. This focus, while crucial for understanding harmful online communications, inadequately addresses the broader spectrum of dehumanization. Specifically, it overlooks the subtler forms of dehumanization that, despite not being overtly offensive, still perpetuate harmful biases against marginalized groups in online interactions. These subtler forms can insidiously reinforce negative stereotypes and biases without explicit offensiveness, making them harder to detect yet equally damaging. Recognizing this gap, we use different sampling methods to collect a theory-informed bilingual dataset from Twitter and Reddit. Using crowdworkers and experts to annotate 16,000 instances on a document- and span-level, we show that our dataset covers the different dimensions of dehumanization. This dataset serves as both a training resource for machine learning models and a benchmark for evaluating future dehumanization detection techniques. To demonstrate its effectiveness, we fine-tune ML models on this dataset, achieving performance that surpasses state-of-the-art models in zero and few-shot in-context settings.

Beyond the Explicit: A Bilingual Dataset for Dehumanization Detection in Social Media

TL;DR

This work tackles the gap in dehumanization detection by introducing a large bilingual (English–German) dataset drawn from and Reddit, containing over 490,000 candidate items and 16,000 annotated instances across animalistic/mechanistic and explicit/implicit dimensions. Using theory-informed sampling, expert crowd annotation, and span-level labeling, the authors demonstrate substantial performance gains for 14 modeling approaches, with notable improvements in zero- and few-shot settings when fine-tuned on the new data. The study also shows robust cross-lingual generalization and provides a detailed error analysis that informs future model improvements. By publishing the full candidate pool and associated predictions, the work offers a valuable resource for training, evaluation, and ongoing data collection in dehumanization detection with practical implications for online safety and bias mitigation.

Abstract

Digital dehumanization, although a critical issue, remains largely overlooked within the field of computational linguistics and Natural Language Processing. The prevailing approach in current research concentrating primarily on a single aspect of dehumanization that identifies overtly negative statements as its core marker. This focus, while crucial for understanding harmful online communications, inadequately addresses the broader spectrum of dehumanization. Specifically, it overlooks the subtler forms of dehumanization that, despite not being overtly offensive, still perpetuate harmful biases against marginalized groups in online interactions. These subtler forms can insidiously reinforce negative stereotypes and biases without explicit offensiveness, making them harder to detect yet equally damaging. Recognizing this gap, we use different sampling methods to collect a theory-informed bilingual dataset from Twitter and Reddit. Using crowdworkers and experts to annotate 16,000 instances on a document- and span-level, we show that our dataset covers the different dimensions of dehumanization. This dataset serves as both a training resource for machine learning models and a benchmark for evaluating future dehumanization detection techniques. To demonstrate its effectiveness, we fine-tune ML models on this dataset, achieving performance that surpasses state-of-the-art models in zero and few-shot in-context settings.
Paper Structure (51 sections, 1 equation, 10 figures, 9 tables)

This paper contains 51 sections, 1 equation, 10 figures, 9 tables.

Figures (10)

  • Figure 1: Conceptualization of Dehumanization based on haslam-reviewbain-haslam-book
  • Figure 2: Multi-step pipeline for selecting potential dehumanizing instances for annotation. It includes data filtering based on predefined targets, personal and demonstrative pronouns, three distinct sampling strategies—theory-informed instance selection, retrieval from existing datasets, and LLM-generated instance matching—and the final annotation process via crowd-workers on Prolific.
  • Figure 3: Semantic space of (a) explicit and (b) implicit candidates that are positively labeled by the crowd as dehumanization.
  • Figure 4: Test F1-Scores for the four different dimensions of dehumanization (see color bars) on English Reddit and Twitter data for different model architectures.
  • Figure 5: Cross-lingual analysis of explicit dehumanization models. The labels on top indicate the evaluation language. Cross-lingual models (hatched bars) were fine-tuned on the opposite language.
  • ...and 5 more figures