CLJan 21, 2025

Challenges in Expanding Portuguese Resources: A View from Open Information Extraction

Marlo Souza, Bruno Cabral, Daniela Claro, Lais Salvador

arXiv:2501.11851v12.7h-index: 3

Originality Synthesis-oriented

AI Analysis

This addresses the problem of limited datasets for non-English languages in Open IE, enabling development and evaluation for Portuguese, though it is incremental as it extends existing methods to a new language.

The authors tackled the lack of Portuguese resources for Open Information Extraction by creating a high-quality manually annotated corpus, which they validated by evaluating state-of-the-art systems.

Open Information Extraction (Open IE) is the task of extracting structured information from textual documents, independent of domain. While traditional Open IE methods were based on unsupervised approaches, recently, with the emergence of robust annotated datasets, new data-based approaches have been developed to achieve better results. These innovations, however, have focused mainly on the English language due to a lack of datasets and the difficulty of constructing such resources for other languages. In this work, we present a high-quality manually annotated corpus for Open Information Extraction in the Portuguese language, based on a rigorous methodology grounded in established semantic theories. We discuss the challenges encountered in the annotation process, propose a set of structural and contextual annotation rules, and validate our corpus by evaluating the performance of state-of-the-art Open IE systems. Our resource addresses the lack of datasets for Open IE in Portuguese and can support the development and evaluation of new methods and systems in this area.

View on arXiv PDF

Similar