Scene Text Detection


The increased usage of contemporary social networks in conjunction with the high-performance smart phones with built-in cameras has sparked a tremendous increase of available images. Textual information in images constitutes a very rich source of high-level semantics for retrieval and indexing. Although research on document image processing has reached a satisfactory level of success, text detection on natural scenes is still a hard task to solve.


Text Detection in Natural Images with Cellular Automata

A new approach is proposed using Cellular Automata (CA) which strives towards identifying scene text on natural images. Initially, a binary edge map is calculated. Then, taking advantage of the CA flexibility, the transition rules are changing and are applied in four consecutive steps resulting in four time steps CA evolution. Finally, a post-processing technique based on edge projection analysis is employed for high density edge images concerning the elimination of possible false positives. Evaluation results indicate considerable performance gains without sacrificing text detection accuracy.


Text Detection in Natural Images Using Bio-inspired Models

Usually, the text that lies in natural images is composed from both states: dark stimuli over bright background and bright stimuli over dark background. Thus, it can be considered as a visual signal, which stimulates both the OFF-center and ON-center ganglion cells, depending on the stimuli spread over the background.

The proposed method addresses the scene text detection problem by modelling the ON and OFF ganglion center-surround cells that reside on the retina. It uses a region and heuristics-based algorithm to support the emulated center-surround cells.


  • Zagoris, K. and Pratikakis, I., “Text detection on natural images using mnemonic cellular automata”, Journal of Cellular Automata, 9 (2-3), (183-194), 2014
  • K. Zagoris and I. Pratikakis, “Text Detection in Natural Images Using Bio-inspired Models”, Document Analysis and Recognition (ICDAR), 2013 12th International Conference on, (1370-1374), 2013
  • Zagoris, K. and Pratikakis, I., “Scene Text Detection on Images Using Cellular Automata”, 10th International Conference on Cellular Automata for Research and Industry (ACRI'12), 2012

Word Spotting in Historical Handwritten Document Images

Word spotting strategies employed in historical handwritten documents face many challenges due to variation in the writing style and intense degradation. The VCG created a new method that permits effective word spotting in handwritten documents that it relies upon document-oriented local features which take into account information around representative keypoints as well as a matching process that incorporates spatial context in a local proximity search without using any training data. Experimental results on four historical handwritten datasets for two different scenarios (segmentation-based and segmentation-free) using standard evaluation measures show the improved performance achieved by the proposed methodology.

The main novelties of the proposed method are:

  • Use of local features that takes into consideration the handwritten documents particularities. Therefore, it is able to detect meaningful points of the characters that reside in the documents independently of its scaling.
  • It provides consistency between different handwritten writing variations.
  • Use of the same operational pipeline in both segmentation-based and segmentation-free scenarios
  • Incorporation of spatial context in the local search of the matching process by integrating a near neighbor search procedure relative to each keypoint.

The keypoint detection method is able to detect meaningful points of the characters that reside in the documents independently of its scaling. Moreover, the linear quantization and the resulting CCs represent chunks of strokes that correspond to different writing directions between them. A subset of these CCs should be stable between different scaling and handwriting styles as some of those chunks remain the same.

The word matching method is motivated by the Nearest Neighbor Search (NNS) by incorporating a spatial context suitable for document images. The advantage of the proposed matching is three-fold: (i) it enables a local search instead of searching in a brute force manner, (ii) it incorporates spatial context and (iii) it is suitable under both segmentation-based or segmentation-free scenarios.


A Framework for Efficient Transcription of Historical Documents Using Keyword Spotting

The VCG proposed a framework that employs KeyWord Spotting to enhance the efficiency in the manual transcription process, thus, reducing drastically the cost of training data creation. The core principle relies upon the ability of robust document-specific descriptors to produce meaningful similarities between a chosen word image for transcription and the corresponding word images in the full dataset under consideration. In the proposed framework, KWS is coupled with a relevance feedback mechanism which further enhances retrieval performance while being independent to the chosen KWS algorithm. The efficiency of the proposed pipeline is showcased via a user-friendly web-based prototype http://vc.ee.duth.gr/ws/.

The major achievement of the proposed framework is the reduction in time expenses required to achieve transcription data which could feed a Handwriting Text Recognition engine for training. Furthermore, the keyword spotting pipeline is coupled with a relevance feedback mechanism which introduces the user in the retrieval loop, thus, improving the final retrieval performance.

  • K. Zagoris, I. Pratikakis, B. Gatos, “Unsupervised Word Spotting in Historical Handwritten Document Images using Document-oriented Local Features,” in IEEE Transactions on Image Processing, vol.PP, no.99, pp.1-1
  • K. Zagoris, I. Pratikakis, and B. Gatos, “Segmentation-based historical handwritten word spotting using document-specific local features,” in Frontiers in Handwriting Recognition (ICFHR), 2014 14th International Conference on, Sept 2014, pp. 9–14.
  • K. Zagoris, I. Pratikakis, and B. Gatos, “A framework for efficient transcription of historical documents using keyword spotting,” in Historical Document Imaging and Processing (HIP'15), 3rd International Workshop on, August 2015, pp. 9–14.