AIDec 15, 2023

A Novel Dataset for Financial Education Text Simplification in Spanish

Nelson Perez-Rojas, Saul Calderon-Ramirez, Martin Solis-Salazar, Mario Romero-Sandoval, Monica Arias-Monge, Horacio Saggion

arXiv:2312.09897v13.91 citationsh-index: 3

Originality Synthesis-oriented

AI Analysis

This addresses the problem of limited resources for text simplification in Spanish, particularly for visually impaired speakers, but it is incremental as it primarily introduces a new dataset.

The researchers tackled the lack of Spanish datasets for text simplification by creating a new dataset with 5,314 complex-simplified sentence pairs focused on financial education, and they compared it with outputs from GPT-3, Tuner, and MT5 to assess data augmentation feasibility.

Text simplification, crucial in natural language processing, aims to make texts more comprehensible, particularly for specific groups like visually impaired Spanish speakers, a less-represented language in this field. In Spanish, there are few datasets that can be used to create text simplification systems. Our research has the primary objective to develop a Spanish financial text simplification dataset. We created a dataset with 5,314 complex and simplified sentence pairs using established simplification rules. We also compared our dataset with the simplifications generated from GPT-3, Tuner, and MT5, in order to evaluate the feasibility of data augmentation using these systems. In this manuscript we present the characteristics of our dataset and the findings of the comparisons with other systems. The dataset is available at Hugging face, saul1917/FEINA.

View on arXiv PDF

Similar