CL LGFeb 7, 2022

Universal Spam Detection using Transfer Learning of BERT Model

arXiv:2202.03480v13.766 citations

Originality Synthesis-oriented

AI Analysis

This work addresses spam detection across multiple datasets, but it is incremental as it applies an existing method to new data combinations.

The paper tackled universal spam detection by fine-tuning a pre-trained BERT model on four datasets, achieving an overall accuracy of 97% and an F1 score of 0.96.

Deep learning transformer models become important by training on text data based on self-attention mechanisms. This manuscript demonstrated a novel universal spam detection model using pre-trained Google's Bidirectional Encoder Representations from Transformers (BERT) base uncased models with four datasets by efficiently classifying ham or spam emails in real-time scenarios. Different methods for Enron, Spamassain, Lingspam, and Spamtext message classification datasets, were used to train models individually in which a single model was obtained with acceptable performance on four datasets. The Universal Spam Detection Model (USDM) was trained with four datasets and leveraged hyperparameters from each model. The combined model was finetuned with the same hyperparameters from these four models separately. When each model using its corresponding dataset, an F1-score is at and above 0.9 in individual models. An overall accuracy reached 97%, with an F1 score of 0.96. Research results and implications were discussed.

View on arXiv PDF

Similar