9.0CLJul 6
SalAngaBhava: A Sinhala Market Dataset for Aspect-based Sentiment AnalysisLakshani Galwatta, Nisansa de Silva, Sarangi Aththanayake et al.
Sentiment analysis has been a primary domain under Natural Language Processing (NLP) from its inception as it plays a vital role in both real-world and research applications. In high-resource languages, this has been extended a step further, and instead of predicting sentiment at the sentence level, models have been developed to detect more fine-grained sentiments at aspect level. However, in order to conduct this fine-grained Aspect-based Sentiment Analysis (ABSA), datasets annotated with aspects and sentiments toward the said aspects is required. Such datasets are lacking for low-resources languages among which, we can count Sinhala, an Indo-Aryan languages used primarily in Sri Lanka. In this work, we introduce, SalAngaBhava, a new Sinhala Aspect-based Sentiment Analysis dataset which contains Sinhala product reviews that are manually labeled with aspect terms and the associated sentiments (positive, negative, neutral). The data was collected from domain-relevant sources such as user-generated reviews and comments, and was annotated following carefully defined guidelines to ensure consistency and quality. The dataset consists of sentences and aspect-sentiment pairs, encompassing a considerable range of aspects from several domains. The analysis confirms that the dataset is well-structured and sufficiently balanced for ABSA research. This dataset can be used as a benchmark and facilitates further studies related to Sinhala natural language processing, and low-resource sentiment analysis tasks.
1.7LGJul 5
Environmental Drivers of Respiratory Disease: A District Level AnalysisRahim Iqbal, Asfi Ahamed, Izzath Nisfer et al.
Sri Lanka has experienced a decade of progressive forest degradation and rising atmospheric pollution, yet district-level respiratory admissions have paradoxically declined, pointing to the confounding role of healthcare access. This study addresses that gap by constructing an 11-year (2014-2024) panel dataset across all 25 administrative districts, integrating satellite-derived vegetation indices, fire radiative power, pollutant concentrations (particulate matter (PM2.5), nitrogen dioxide (NO2), sulfur dioxide (SO2)), carbon flux metrics and population-normalized respiratory admission rates. Two temporally validated XGBoost models were created for annual district-level respiratory rate (R^2 = 0.937) and monthly PM2.5 concentration (R^2 = 0.976) with generalization validated in 21 out of 25 districts (Mean Absolute Percentage Error (MAPE) <= 20%). Shapley Additive Explanations (SHAP) analysis established that cumulative air quality burden is the overwhelming driver of respiratory rate variance (80.1%), ahead of forest degradation (15.6%) and fire activity (4.3%). The Forest-Air-Health (FAH) Risk Index used these SHAP-derived weights to find the districts with the highest risk: Colombo (FAH = 0.802), Gampaha (0.708), and Kalutara (0.682). These findings present the inaugural evidence-based, district-level framework correlating environmental degradation with respiratory health in Sri Lanka, establishing a quantitative basis for focused public health and environmental policy.