Unconstrained Language Identification Using AShape Codebook

TitleUnconstrained Language Identification Using AShape Codebook
Publication TypeConference Papers
Year of Publication2008
AuthorsZhu G, Yu X, Li Y, Doermann D
Conference NameThe 11th International Conference on Frontiers in Handwritting Recognition (ICFHR 2008)
Date Published2008///
Conference LocationMontreal, Canada
Abstract

We propose a novel approach to language identification in document images containing handwriting and machine printed text using image descriptors constructed from a codebook of shape features. We encode local text structures using scale and rotation invariant codewords, each representing a characteristic shape feature that is generic enough to appear repeatably. We learn a concise, structurally indexed shape codebook from training data by clustering similar features and partitioning the feature space by graph cuts. Our approach is segmentation free and easily extensible. We quantitatively evaluate our approach using a large real-world document image collection, which consists of more than 1,500 documents in 8 languages (Arabic, Chinese, English, Hindi, Japanese, Korean, Russian, and Thai) and contains a complex mixture of handwritten and machine printed content. Experimental results demonstrate the robustness and flexibility of our approach, and show exceptional language identification performance that exceeds the state of art.