CVJun 29, 2021

SDL: New data generation tools for full-level annotated document layout

arXiv:2106.15117v1Has Code
Originality Synthesis-oriented
AI Analysis

This tool addresses the need for annotated document data in low-resource languages, though it is incremental as it builds on existing data generation methods.

The paper tackles the problem of generating annotated document layout data by introducing a tool that provides full-level visual information from character to paragraph positions, and it includes a dataset of 320,000 Vietnamese synthetic document images with instructions for generating similar datasets in other languages.

We present a novel data generation tool for document processing. The tool focuses on providing a maximal level of visual information in a normal type document, ranging from character position to paragraph-level position. It also enables working with a large dataset on low-resource languages as well as providing a mean of processing thorough full-level information of the documented text. The data generation tools come with a dataset of 320000 Vietnamese synthetic document images and an instruction to generate a dataset of similar size in other languages. The repository can be found at: https://github.com/tson1997/SDL-Document-Image-Generation

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes