CVJul 1

Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs

arXiv:2607.005963.8
Predicted impact top 84% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers working on historical document analysis in under-resourced languages, this provides a practical bootstrapping strategy for rapid annotation.

This paper tackles reading order reconstruction in historical Armenian newspapers with complex layouts and limited resources. Their hybrid method combining semantic zone detection with a generative LLM reduces ordering errors by up to 76% over the strongest geometric baseline.

This paper addresses reading order reconstruction in historical Armenian newspapers, which combine complex layouts with limited language resources. We introduce a new annotated dataset of 66 pages and compare geometric heuristics, YOLO-based layout parsing, an end-to-end document model ECLAIR, and a hybrid method combining semantic zone detection with a generative LLM. Our hybrid method achieves the lowest error rates of all evaluated approaches, reducing ordering errors by up to 76% over the strongest geometric baseline, and remains robust in multi-page settings and under noisy OCR. Rather than targeting production the method is designed as a data bootstrapping strategy enabling rapid annotation in highly under-resourced scenarios. Alongside the dataset, we release a specialized Tesseract OCR model for historical Armenian print.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes