Framework for Performance Evaluation of Face, Text, and Vehicle Detection and Tracking in Video: Data, Metrics, and Protocol

Top Cited Papers

31 March 2008

journal article
research article
Published by Institute of Electrical and Electronics Engineers (IEEE) in IEEE Transactions on Pattern Analysis and Machine Intelligence

Vol. 31 (2), 319-336
https://doi.org/10.1109/tpami.2008.57

Abstract

Common benchmark data sets, standardized performance metrics, and baseline algorithms have demonstrated considerable impact on research and development in a variety of application domains. These resources provide both consumers and developers of technology with a common framework to objectively compare the performance of different algorithms and algorithmic improvements. In this paper, we present such a framework for evaluating object detection and tracking in video: specifically for face, text, and vehicle objects. This framework includes the source video data, ground-truth annotations (along with guidelines for annotation), performance metrics, evaluation protocols, and tools including scoring software and baseline algorithms. For each detection and tracking task and supported domain, we developed a 50-clip training set and a 50-clip test set. Each data clip is approximately 2.5 minutes long and has been completely spatially/temporally annotated at the I-frame level. Each task/domain, therefore, has an associated annotated corpus of approximately 450,000 frames. The scope of such annotation is unprecedented and was designed to begin to support the necessary quantities of data for robust machine learning approaches, as well as a statistically significant comparison of the performance of algorithms. The goal of this work was to systematically address the challenges of object detection and tracking through a common evaluation framework that permits a meaningful objective comparison of techniques, provides the research community with sufficient data for the exploration of automatic modeling techniques, encourages the incorporation of objective evaluation into the development process, and contributes useful lasting resources of a scale and magnitude that will prove to be extremely useful to the computer vision research community for years to come.

Keywords

This publication has 28 references indexed in Scilit:

On-road vehicle detection: a review
IEEE Transactions on Pattern Analysis and Machine Intelligence, 2006
PETS Metrics: On-Line Performance Evaluation Service
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2006
Performance evaluation of fingerprint verification systems
IEEE Transactions on Pattern Analysis and Machine Intelligence, 2005
Overview of the Face Recognition Grand Challenge
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2005
Text information extraction in images and video: a survey
Pattern Recognition, 2004
Extraction of special effects caption text events from digital video
International Journal on Document Analysis and Recognition (IJDAR), 2003
Real-time detection of moving vehicles
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2003
Fast vehicle detection with probabilistic feature grouping and its application to vehicle tracking
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2003
The FERET evaluation methodology for face-recognition algorithms
IEEE Transactions on Pattern Analysis and Machine Intelligence, 2000
VISATRAM: a real-time vision system for automatic traffic monitoring
Image and Vision Computing, 2000

Cited by 369 articles