CLJun 24

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar

arXiv:2606.2601523.6Has Code
Predicted impact top 24% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners working on text detoxification in low-resource languages, this work provides a new SOTA system and dataset for Tatar, though it is incremental in nature.

The paper presents Tatoxa, a state-of-the-art system for text detoxification in the low-resource Tatar language, which outperforms existing open-source and proprietary LLMs on key quality metrics. It also introduces a new Tatar detoxification dataset and shows that cross-lingual transfer from Russian performs worse than training on native Tatar data.

Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring the safety of online communities and protecting users. However, low resource languages such as Tatar have received little research attention. In this paper we present Tatoxa, a novel state-of-the-art system for text detoxification in the Tatar language. Comparative experiments show that the proposed approach outperforms existing open source and proprietary commercial LLMs on key quality metrics. We also introduce a new dataset for text detoxification in Tatar, designed for fine tuning and evaluation in low resource settings. Finally, cross lingual transfer experiments indicate that transfer from other languages, including the culturally close Russian, performs significantly worse than training on native Tatar data even when a large Russian corpus is available.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes