Semantic-Guided Reading Order Reconstruction in Historical Armenian Newspapers with LLMs
For researchers working on historical document analysis in under-resourced languages, this provides a practical bootstrapping strategy for rapid annotation.
This paper tackles reading order reconstruction in historical Armenian newspapers with complex layouts and limited resources. Their hybrid method combining semantic zone detection with a generative LLM reduces ordering errors by up to 76% over the strongest geometric baseline.
This paper addresses reading order reconstruction in historical Armenian newspapers, which combine complex layouts with limited language resources. We introduce a new annotated dataset of 66 pages and compare geometric heuristics, YOLO-based layout parsing, an end-to-end document model ECLAIR, and a hybrid method combining semantic zone detection with a generative LLM. Our hybrid method achieves the lowest error rates of all evaluated approaches, reducing ordering errors by up to 76% over the strongest geometric baseline, and remains robust in multi-page settings and under noisy OCR. Rather than targeting production the method is designed as a data bootstrapping strategy enabling rapid annotation in highly under-resourced scenarios. Alongside the dataset, we release a specialized Tesseract OCR model for historical Armenian print.