1,720,992 research outputs found
Recognition and tracking of embedded text in sport videos: a case study
Abstract - We present an athlete recognition module designed for broadcast videos, forming part of a system designed for the personalization of sport video broadcast. The aim of this module is the identification of athletes in the scene through the reading of names or numbers printed on their uniforms and to identify frames where athletes are visible. Using an adaptation of a previously published algorithm we extract text from individual frames and then read these candidates by means of an optical character recognizer (OCR). The OCR-ed text is then compared to an a-priori list of athletes' names (or numbers), to provide a presence score for each athlete. Text regions are subsequently tracked in following frames using a template matching technique. In this way blurred or distorted text, normally unreadable by the OCR, can also be exploited to provide a denser labelling of the video sequences
Low Level Processing of Color Document Images
Color is an extremely useful cue that is used for adding information in documents. Document image understanding takes advantage from the implicit or explicit knowledge on the use of color in document production. This knowledge has to be exploited starting from the low level tasks. In this report low level image processing functions, particularly suited for the processing of scanned color documents images, are considered and analyzed in a brief survey. Extensions of some edge detection and edge preserving smoothing algorithms to the color domain are introduce
Extraction of Polygonal Frames from Color Documents for Page Decomposition
Graphic accents are often used in the design of complex documents in order to emphasize particular information. Words or illustrations are surrounded by a border line or highlighted by means of a colored background. This paper presents a method for the
automatic extraction of document layout items, called {\em frames}, having polygonal shape and/or a uniformly colored background. As frames break the normal text flow, frame detection is a fundamental step of the document layout analysis in a document understanding system.
The presented method relies on a color region growing algorithm and on straight edges extractor. The shape analysis of the obtained regions permits to localize the frames with their attributes. In order to reduce computation time and to return only specific patterns, the method exploits information about a model of the frames to be detected
such as shape, skew or size, possibly supplied by the user or depending on the specific document class.
The presented algorithm is assessed on a page databases containing more than 675 framed items. The evaluation is based on a novel tree matching method that takes into account the frame hierarchy and their shap
Automatic Identification and Skew Estimation of Text Lines in Real Scene Images
The paper presents an algorithm for the automatic localization of text embedded in complex images. It detects the spatial position and the skew of the text lines which are present in the scene and returns a binary representation of each text line. For comparison reasons we report a brief review of recent works in this field. Strength of the presented algorithm are independence of text skew and of presence of connected text. Two main assumptions are made about text lines: to have a straight base line, to be composed of characters with comparable heights. After a pre-processing step the input image is segmented in order to obtain a set of connected components which represent the basic elements of the algorithm. Several heuristics are proposed to characterized text components with respect to non-text ones: they depend both on the geometrical features of single components and on the geometrical and spatial relations among components. According to these heuristics several components are discarded and the retained ones are grouped into text lines candidates by means of a divisive hierarchical clustering procedure. In the experimental session we describe the application of the algorithm to the extraction of text lines from images of 100 book covers. Results about skew estimation are also reported
Boosting Fisher vector based scoring functions for person re-identification
In recent years, much effort has been put into the development of novel algorithms to solve the person re-identification problem. The goal is to match a given person's image against a gallery of people. In this paper, we propose a single-shot supervised method to compute a scoring function that, when applied to a pair of images, provides a score expressing the likelihood that they depict the same individual. The method is characterized by: (i) the usage of a set of local image descriptors based on Fisher vectors, (ii) the training of a pool of scoring functions based on the local descriptors, and (iii) the construction of a strong scoring function by means of an adaptive boosting procedure. The method has been tested on four data-sets and results have been compared with state-of-the-art methods clearly showing superior performance
VCB: A Video Clip Browser and More
Browsing is an interactive process through which the user decide when and where to focus the attention within a large database, by example in order to retrieve information in a video stream. This paper presents the 'Video Clip Browser', a set of tools designed to easily navigate into video clips, as well as some utilities for the automatic search of particular frames of the video stream. A peculiarity of this system is that the browsing tools can be used in an interleaved way giving high flexibility to the navigation proces
The Indexing Function in Document Management Systems
In this note we analyze the indexing function of commercial electronic document management systems. In particular, we have focused our attention to the relation between indexing and the use of an OCR module in archiving products. On the basis of this relation, we have distinguished products into three classes: (1) indexing by keywords wighout using OCR; (2) field-based indexing where OCR is aplied only to user defined fields; (3) full-text indexing, where OCR is applied to the whole document or to its textual regions. As a conclusive remark, we highlight that no one of the analyzed product provides the capability of automatic segmentation and logical labeling of the document: a necessary feature for development of intelligent document indexing and retrieval system
Progetto CODICE*: Alcune considerazioni sulle modalità di valutazione del sistema
Queste note raccolgono alcune considerazioni relative alla definizione delle modalità di valutazione del sistema di analisi ed interpretazione di documenti complessi in via di sviluppo nell’ambito del progetto CODICE. Viene messo in evidenza che la possibilità di definire metodi di valutazione generali è attuabile solo per alcune sottoparti del sistema se non per specifiche procedure, mentre la definizione di una validazione globale richiede che sia fissato lo scenario applicativ
Scene Text Recognition and Tracking to Identify Athletes in Sport Videos
We present an athlete identification module forming part of a system for the personalization of sport video broadcasts. The aim of this module is the localization of athletes in the scene, their identification through the reading of names or numbers printed on their uniforms, and the labelling of frames where
athletes are visible. Building upon a previously published algorithm we extract text from individual frames and read these candidates by means of an optical character recognizer (OCR). The OCR-ed text is then compared to a known list of athletes’ names (or numbers), to provide a presence score for each athlete. Text regions are tracked in subsequent frames using a template matching technique. In this way blurred or distorted text, normally unreadable by the OCR, is exploited to provide a denser labelling of the video sequences. Extensive experiments show that themethod proposed is fast, robust and reliable, out-performing results of other systems in the literature
Context Driven Text Segmentation and Recognition
A segmentation algorithm is presented within a text recognition system. The output of an OCR system and the contextual analysis based on the dictionary method are used jointly to segment a line of text. A distance definition is introduced between a pattern representing a bitmap with the input text line and a string representing an entry of the dictionary. The algorithm is used in a system that aims at recognizing a book by reading the scripts on its cover and by comparing the recognized text with a stored database. Preliminary experiments are presented for both the line recognition and book recognition task
- …
