CLAILGMay 15, 2024

Bridging the gap in online hate speech detection: a comparative analysis of BERT and traditional models for homophobic content identification on X/Twitter

arXiv:2405.09221v18 citationsh-index: 3Has CodeAppl Comput Eng
Originality Incremental advance
AI Analysis

This addresses a gap in online safety and inclusivity for LGBTQIA+ communities by improving detection of homophobic hate speech, though it is incremental as it builds on existing sentiment analysis methods.

The study tackled the underrepresentation of homophobia in online hate speech detection by comparing BERT and traditional models for identifying homophobic content on X/Twitter, finding that BERT outperforms traditional methods but performance depends on validation techniques, and releasing the largest open-source labelled English dataset for this purpose.

Our study addresses a significant gap in online hate speech detection research by focusing on homophobia, an area often neglected in sentiment analysis research. Utilising advanced sentiment analysis models, particularly BERT, and traditional machine learning methods, we developed a nuanced approach to identify homophobic content on X/Twitter. This research is pivotal due to the persistent underrepresentation of homophobia in detection models. Our findings reveal that while BERT outperforms traditional methods, the choice of validation technique can impact model performance. This underscores the importance of contextual understanding in detecting nuanced hate speech. By releasing the largest open-source labelled English dataset for homophobia detection known to us, an analysis of various models' performance and our strongest BERT-based model, we aim to enhance online safety and inclusivity. Future work will extend to broader LGBTQIA+ hate speech detection, addressing the challenges of sourcing diverse datasets. Through this endeavour, we contribute to the larger effort against online hate, advocating for a more inclusive digital landscape. Our study not only offers insights into the effective detection of homophobic content by improving on previous research results, but it also lays groundwork for future advancements in hate speech analysis.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes