CLOct 1, 2020

"Did you really mean what you said?" : Sarcasm Detection in Hindi-English Code-Mixed Data using Bilingual Word Embeddings

Akshita Aggarwal, Anshul Wadhawan, Anshima Chaudhary, Kavita Maurya

arXiv:2010.00310v3994 citations

Originality Synthesis-oriented

AI Analysis

This addresses sarcasm detection for social media users in multilingual contexts, but it is incremental as it applies existing methods to a new dataset.

The paper tackled sarcasm detection in Hindi-English code-mixed tweets by proposing a deep learning approach using bilingual word embeddings, achieving a state-of-the-art accuracy of 78.49% with attention-based Bi-directional LSTMs.

With the increased use of social media platforms by people across the world, many new interesting NLP problems have come into existence. One such being the detection of sarcasm in the social media texts. We present a corpus of tweets for training custom word embeddings and a Hinglish dataset labelled for sarcasm detection. We propose a deep learning based approach to address the issue of sarcasm detection in Hindi-English code mixed tweets using bilingual word embeddings derived from FastText and Word2Vec approaches. We experimented with various deep learning models, including CNNs, LSTMs, Bi-directional LSTMs (with and without attention). We were able to outperform all state-of-the-art performances with our deep learning models, with attention based Bi-directional LSTMs giving the best performance exhibiting an accuracy of 78.49%.

View on arXiv PDF

Similar