Table of Contents
Fetching ...

A Big Data-empowered System for Real-time Detection of Regional Discriminatory Comments on Vietnamese Social Media

An Nghiep Huynh, Thanh Dat Do, Trong Hop Do

TL;DR

This work has built the ViRDC (Vietnamese Regional Discrimination Comments) dataset, which contains comments from social media platforms, providing a valuable resource for further research and development and developed the system on the Apache Spark framework to efficiently handle increasing data inputs during streaming.

Abstract

Regional discrimination is a persistent social issue in Vietnam. While existing research has explored hate speech in the Vietnamese language, the specific issue of regional discrimination remains under-addressed. Previous studies primarily focused on model development without considering practical system implementation. In this work, we propose a task called Detection of Regional Discriminatory Comments on Vietnamese Social Media, leveraging the power of machine learning and transfer learning models. We have built the ViRDC (Vietnamese Regional Discrimination Comments) dataset, which contains comments from social media platforms, providing a valuable resource for further research and development. Our approach integrates streaming capabilities to process real-time data from social media networks, ensuring the system's scalability and responsiveness. We developed the system on the Apache Spark framework to efficiently handle increasing data inputs during streaming. Our system offers a comprehensive solution for the real-time detection of regional discrimination in Vietnam.

A Big Data-empowered System for Real-time Detection of Regional Discriminatory Comments on Vietnamese Social Media

TL;DR

This work has built the ViRDC (Vietnamese Regional Discrimination Comments) dataset, which contains comments from social media platforms, providing a valuable resource for further research and development and developed the system on the Apache Spark framework to efficiently handle increasing data inputs during streaming.

Abstract

Regional discrimination is a persistent social issue in Vietnam. While existing research has explored hate speech in the Vietnamese language, the specific issue of regional discrimination remains under-addressed. Previous studies primarily focused on model development without considering practical system implementation. In this work, we propose a task called Detection of Regional Discriminatory Comments on Vietnamese Social Media, leveraging the power of machine learning and transfer learning models. We have built the ViRDC (Vietnamese Regional Discrimination Comments) dataset, which contains comments from social media platforms, providing a valuable resource for further research and development. Our approach integrates streaming capabilities to process real-time data from social media networks, ensuring the system's scalability and responsiveness. We developed the system on the Apache Spark framework to efficiently handle increasing data inputs during streaming. Our system offers a comprehensive solution for the real-time detection of regional discrimination in Vietnam.

Paper Structure

This paper contains 29 sections, 7 figures, 6 tables.

Figures (7)

  • Figure 1: Data Labeling Process
  • Figure 2: Preprocessing Steps
  • Figure 3: Label Distribution in all Files
  • Figure 4: Comment Length Distribution
  • Figure 5: The architecture of the proposed system
  • ...and 2 more figures