CLAIJun 30

Hate Speech Detection in Turkish and Arabic Languages: A Comprehensive Study

arXiv:2607.0014311.3
Predicted impact top 72% in CL · last 90 daysOriginality Incremental advance
AI Analysis

This work provides resources and models for hate speech detection in under-resourced languages (Turkish and Arabic), addressing a societal need for content moderation.

The paper introduces a comprehensive hate speech dataset in Turkish and Arabic covering multiple topics and develops BERT-based models for hate category classification, intensity prediction, target identification, and span detection, achieving state-of-the-art results.

Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings, and ethnic cleansing. Societies grappling with this issue, particularly when hate speech targets specific groups based on religion, race, ethnicity, culture, nationality, or migration status, face the challenge of balancing freedom of expression with the need for effective content moderation on widely used online platforms. In response to this challenge, we introduce a comprehensive hate speech dataset covering five distinct topics in Turkish: refugees, the Israel-Palestine conflict, anti-Greek sentiment in Turkey, ethnic or religious communities (Alevis, Armenians, Arabs, Jews, and Kurds), and LGBTI+, alongside one topic in Arabic (refugees). In addition, we develop state-of-the-art BERT-based models to address multiple dimensions of hate speech analysis, including hate category classification, hate intensity prediction, target identification, and hate speech span detection, enabling a comprehensive understanding of hateful content in online discourse.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes