A Unified Framework for Tracking Based Text Detection and Recognition from Web Videos

Abstract

Video text extraction plays an important role for multimedia understanding and retrieval. Most previous research efforts are conducted within individual frames. A few of recent methods, which pay attention to text tracking using multiple frames, however, do not effectively mine the relations among text detection, tracking and recognition. In this paper, we propose a generic Bayesian-based framework of Tracking based Text Detection And Recognition (T2DAR) from web videos for embedded captions, which is composed of three major components, i.e., text tracking, tracking based text detection, and tracking based text recognition. In this unified framework, text tracking is first conducted by tracking-by-detection. Tracking trajectories are then revised and refined with detection or recognition results. Text detection or recognition is finally improved with multi-frame integration. Moreover, a challenging video text (embedded caption text) database (USTB-VidTEXT) is constructed and publicly available. A variety of experiments on this dataset verify that our proposed approach largely improves the performance of text detection and recognition from web videos.

Funding Information

National Natural Science Foundation of China (61473036)

This publication has 42 references indexed in Scilit:

Robust Text Detection in Natural Scene Images
IEEE Transactions on Pattern Analysis and Machine Intelligence, 2013
The Hungarian Method for the Assignment Problem
Published by Springer Science and Business Media LLC ,2009
Semi-supervised On-Line Boosting for Robust Tracking
Lecture Notes in Computer Science, 2008
Object count/area graphs for the evaluation of object detection and segmentation algorithms
International Journal on Document Analysis and Recognition (IJDAR), 2006
Camera-based analysis of text and documents: a survey
International Journal on Document Analysis and Recognition (IJDAR), 2005
Text information extraction in images and video: a survey
Pattern Recognition, 2004
A Boosted Particle Filter: Multitarget Detection and Tracking
Lecture Notes in Computer Science, 2004
Localizing and segmenting text in images and videos
IEEE Transactions on Circuits and Systems for Video Technology, 2002
Automatic text segmentation and text recognition for video indexing
Multimedia Systems, 2000
Automatic text detection and tracking in digital video
IEEE Transactions on Image Processing, 2000

Cited by 54 articles