CVNov 17, 2014

AlexU-Word: A New Dataset for Isolated-Word Closed-Vocabulary Offline Arabic Handwriting Recognition

arXiv:1411.4670v1
Originality Synthesis-oriented
AI Analysis

This dataset addresses the problem of closed-vocabulary Arabic handwriting recognition for researchers, but it is incremental as it builds on existing data collection efforts.

The authors introduced AlexU-Word, a new dataset for isolated-word offline Arabic handwriting recognition, containing 25,114 samples of 109 unique words from 907 writers, and achieved 92.16% accuracy using a SIFT-based descriptor and ANN.

In this paper, we introduce the first phase of a new dataset for offline Arabic handwriting recognition. The aim is to collect a very large dataset of isolated Arabic words that covers all letters of the alphabet in all possible shapes using a small number of simple words. The end goal is to collect a very large dataset of segmented letter images, which can be used to build and evaluate Arabic handwriting recognition systems that are based on segmented letter recognition. The current version of the dataset contains $25114$ samples of $109$ unique Arabic words that cover all possible shapes of all alphabet letters. The samples were collected from $907$ writers. In its current form, the dataset can be used for the problem of closed-vocabulary word recognition. We evaluated a number of window-based descriptors and classifiers on this task and obtained an accuracy of $92.16\%$ using a SIFT-based descriptor and ANN.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes