Compiling and Processing Historical and Contemporary Portuguese Corpora
This work provides a resource for researchers in linguistics and natural language processing by compiling and processing Portuguese corpora, but it is incremental as it applies existing methods to new data.
The authors tackled the challenge of processing three large Portuguese corpora, including historical and contemporary texts, by developing a framework for pre-processing, segmentation, annotation, indexing, and querying, and they reported on published research papers utilizing these corpora.
This technical report describes the framework used for processing three large Portuguese corpora. Two corpora contain texts from newspapers, one published in Brazil and the other published in Portugal. The third corpus is Colonia, a historical Portuguese collection containing texts written between the 16th and the early 20th century. The report presents pre-processing methods, segmentation, and annotation of the corpora as well as indexing and querying methods. Finally, it presents published research papers using the corpora.