PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction
This work addresses the lack of OCR benchmarks for European Portuguese, providing a resource for evaluating and improving text extraction in this low-resource language.
The authors introduce PorTEXTO, the first benchmark for modern European Portuguese visual text extraction, and find that specialized multilingual data drives performance more than model size or resolution, motivating open OCR resources.
European Portuguese (pt-PT) is largely absent from OCR benchmarks, which skew toward high-resource languages. The few benchmarks that cover pt-PT focus on historical artifacts and literature. This work addresses modern OCR applications, introducing PorTEXTO, the first benchmark for contemporary and culturally relevant pt-PT visual text extraction. To ascertain quality, we employ an annotation pipeline combining transcriptions from a frontier LVLM with exhaustive review by native speakers. We observe a sharp performance drop from synthetic to real world samples in most models, and find that, currently, specialized multilingual data is a better driver for pt-PT performance than model size or resolution budget, motivating the release of open pt-PT OCR resources.