Grounding Toxicity in Real-World Events across Languages
Wondimagegnhue Tsegaye Tufa, Ilia Markov, Piek Vossen
TL;DR
This study investigates how major real-world events drive the origin and spread of toxicity in multilingual online discussions. It analyzes 4.5 million Reddit comments across six languages (Dutch, English, German, Arabic, Turkish, Spanish) tied to 15 events from 2020–2023, using a lexicon-based toxicity detector supplemented by GPT-4 and Perspective API for evaluation. The authors examine toxicity alongside sentiment and emotion via the NRC lexicon, performing both within-language and cross-language analyses to reveal event- and language-specific patterns, including delayed toxicity peaks and strong cross-language relationships. The work provides a multilingual, event-grounded framework for understanding online toxicity and releases the data and code to enable further research and cross-cultural insights.
Abstract
Social media conversations frequently suffer from toxicity, creating significant issues for users, moderators, and entire communities. Events in the real world, like elections or conflicts, can initiate and escalate toxic behavior online. Our study investigates how real-world events influence the origin and spread of toxicity in online discussions across various languages and regions. We gathered Reddit data comprising 4.5 million comments from 31 thousand posts in six different languages (Dutch, English, German, Arabic, Turkish and Spanish). We target fifteen major social and political world events that occurred between 2020 and 2023. We observe significant variations in toxicity, negative sentiment, and emotion expressions across different events and language communities, showing that toxicity is a complex phenomenon in which many different factors interact and still need to be investigated. We will release the data for further research along with our code.
