CVDLJun 14

A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts

arXiv:2606.159872.8
Predicted impact top 92% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

This work provides a challenging benchmark for low-resource HTR, addressing the needs of researchers working on historical documents with rare scripts and degraded conditions.

The authors introduce SCAM, a new line-level dataset for Handwritten Text Recognition from Sahidic Coptic ancient manuscripts, and benchmark several state-of-the-art methods, revealing a significant performance gap compared to well-resourced modern scripts.

In this work, we target Handwritten Text Recognition (HTR) in low-resource scenarios, which arise from underrepresented languages, rare scripts, and degraded visual conditions typical of historical documents. We introduce SCAM (Sahidic Coptic Ancient Manuscripts), a new line-level dataset built from digitized ancient manuscripts written in the extinct Sahidic Coptic dialect. The dataset reflects a realistic and challenging setting, as it combines heterogeneous acquisition conditions across libraries with typical manuscript degradations such as ink fading, bleed-through, and material deterioration. In addition to visual complexity, SCAM poses significant linguistic challenges due to the scarcity of resources for Sahidic Coptic, its uncommon alphabet, and dialect-specific diacritics. To support research in low-resource HTR, we benchmark several state-of-the-art approaches based on different paradigms, highlighting their limitations and strengths in this setting. Our results underline the gap between current HTR performance on well-resourced modern scripts and historically grounded, low-resource scenarios, thus providing a reference point for future developments.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes