CLJan 8, 2020

LTP: A New Active Learning Strategy for CRF-Based Named Entity Recognition

Mingyi Liu, Zhiying Tu, Tong Zhang, Tonghua Su, Zhongjie Wang

arXiv:2001.02524v23.535 citationsHas Code

Originality Incremental advance

AI Analysis

This work addresses annotation efficiency for NLP practitioners, but it is incremental as it builds on existing uncertainty-based active learning approaches.

The paper tackles the high annotation cost in named entity recognition by proposing a new active learning strategy called LTP, which reduces annotation tokens while achieving slightly better accuracy and F1-score than traditional methods.

In recent years, deep learning has achieved great success in many natural language processing tasks including named entity recognition. The shortcoming is that a large amount of manually-annotated data is usually required. Previous studies have demonstrated that active learning could elaborately reduce the cost of data annotation, but there is still plenty of room for improvement. In real applications we found existing uncertainty-based active learning strategies have two shortcomings. Firstly, these strategies prefer to choose long sequence explicitly or implicitly, which increase the annotation burden of annotators. Secondly, some strategies need to invade the model and modify to generate some additional information for sample selection, which will increase the workload of the developer and increase the training/prediction time of the model. In this paper, we first examine traditional active learning strategies in a specific case of BiLstm-CRF that has widely used in named entity recognition on several typical datasets. Then we propose an uncertainty-based active learning strategy called Lowest Token Probability (LTP) which combines the input and output of CRF to select informative instance. LTP is simple and powerful strategy that does not favor long sequences and does not need to invade the model. We test LTP on multiple datasets, and the experiments show that LTP performs slightly better than traditional strategies with obviously less annotation tokens on both sentence-level accuracy and entity-level F1-score. Related code have been release on https://github.com/HIT-ICES/AL-NER

View on arXiv PDF Code

Similar