MultiTACRED: A Multilingual Version of the TAC Relation Extraction DatasetLeonhard Hennig, Philippe Thomas, Sebastian Möller
Relation extraction (RE) is a fundamental task in information extraction, whose extension to multilingual settings has been hindered by the lack of supervised resources comparable in size to large English datasets such as TACRED (Zhang et al., 2017). To address this gap, we introduce the MultiTACRED dataset, covering 12 typologically diverse languages from 9 language families, which is created by machine-translating TACRED instances and automatically projecting their entity annotations. We analyze translation and annotation projection quality, identify error categories, and experimentally evaluate fine-tuned pretrained mono- and multilingual language models in common transfer learning scenarios. Our analyses show that machine translation is a viable strategy to transfer RE instances, with native speakers judging more than 83% of the translated instances to be linguistically and semantically acceptable. We find monolingual RE model performance to be comparable to the English original for many of the target languages, and that multilingual models trained on a combination of English and target language data can outperform their monolingual counterparts. However, we also observe a variety of translation and annotation projection errors, both due to the MT systems and linguistic features of the target languages, such as pronoun-dropping, compounding and inflection, that degrade dataset quality and RE model performance.
32.0CLApr 7, 2020
A German Corpus for Fine-Grained Named Entity Recognition and Relation Extraction of Traffic and Industry EventsMartin Schiersch, Veselina Mironova, Maximilian Schmitt et al.
Monitoring mobility- and industry-relevant events is important in areas such as personal travel planning and supply chain management, but extracting events pertaining to specific companies, transit routes and locations from heterogeneous, high-volume text streams remains a significant challenge. This work describes a corpus of German-language documents which has been annotated with fine-grained geo-entities, such as streets, stops and routes, as well as standard named entity types. It has also been annotated with a set of 15 traffic- and industry-related n-ary relations and events, such as accidents, traffic jams, acquisitions, and strikes. The corpus consists of newswire texts, Twitter messages, and traffic reports from radio stations, police and railway companies. It allows for training and evaluating both named entity recognition algorithms that aim for fine-grained typing of geo-entities, as well as n-ary relation extraction systems.
2.9CLOct 24, 2017
Clickbait Identification using Neural NetworksPhilippe Thomas
This paper presents the results of our participation in the Clickbait Detection Challenge 2017. The system relies on a fusion of neural networks, incorporating different types of available informations. It does not require any linguistic preprocessing, and hence generalizes more easily to new domains and languages. The final combined model achieves a mean squared error of 0.0428, an accuracy of 0.826, and a F1 score of 0.564. According to the official evaluation metric the system ranked 6th of the 13 participating teams.
4.1SEJun 13, 2013
Improving production process performance thanks to neuronal analysisMélanie Noyel, Philippe Thomas, Patrick Charpentier et al.
Product quality level is become a key factor for companies' competitiveness. A lot of time and money are required to ensure and guaranty it. Besides, motivated by the need of traceability, collecting production data is now commonplace in most companies. Our paper aims to show that we can ensure the required quality thanks to an "on-line quality approch" and proposes a neural network based process to determine the optimal setting for production machines. We will illustrate this with the Acta-Mobilier case, which is a high quality lacquerer company.