TY - CONF T1 - Stroke-like Pattern Noise Removal in Binary Document Images T2 - International Conference on Document Analysis and Recognition Y1 - 2011 A1 - Agrawal,Mudit A1 - David Doermann AB - This paper presents a two-phased stroke-like pattern noise (SPN) removal algorithm for binary document images. The proposed approach aims at understanding script-independent prominent text component features using supervised classification as a first step. It then uses their cohesiveness and stroke-width properties to filter and associate smaller text components with them using an unsupervised classification technique. In order to perform text extraction, and hence noise removal, at diacritic-level, this divide-and-conquer technique does not assume the availability of accurate and large amounts of ground-truth data at component-level for training purposes. The method was tested on a collection of degraded and noisy, machine-printed and handwritten binary Arabic text documents. Results show pixel-level precision and recall of 98% and 97% respectively. JA - International Conference on Document Analysis and Recognition ER -