Hierarchical Recurrent Neural Encoder for Video Representation with Application to Captioning
- 1 June 2016
- conference paper
- conference paper
- Published by Institute of Electrical and Electronics Engineers (IEEE)
- p. 1029-1038
- https://doi.org/10.1109/cvpr.2016.117
Abstract
Recently, deep learning approach, especially deep Convolutional Neural Networks (ConvNets), have achieved overwhelming accuracy with fast processing speed for image classification. Incorporating temporal structure with deep ConvNets for video representation becomes a fundamental problem for video content analysis. In this paper, we propose a new approach, namely Hierarchical Recurrent Neural Encoder (HRNE), to exploit temporal information of videos. Compared to recent video representation inference approaches, this paper makes the following three contributions. First, our HRNE is able to efficiently exploit video temporal structure in a longer range by reducing the length of input information flow, and compositing multiple consecutive inputs at a higher level. Second, computation operations are significantly lessened while attaining more non-linearity. Third, HRNE is able to uncover temporal tran-sitions between frame chunks with different granularities, i.e. it can model the temporal transitions between frames as well as the transitions between segments. We apply the new method to video captioning where temporal information plays a crucial role. Experiments demonstrate that our method outperforms the state-of-the-art on video captioning benchmarks.Keywords
This publication has 22 references indexed in Scilit:
- A Hierarchical Recurrent Encoder-Decoder for Generative Context-Aware Query SuggestionPublished by Association for Computing Machinery (ACM) ,2015
- A dataset for Movie DescriptionPublished by Institute of Electrical and Electronics Engineers (IEEE) ,2015
- A discriminative CNN video representation for event detectionPublished by Institute of Electrical and Electronics Engineers (IEEE) ,2015
- Translating Videos to Natural Language Using Deep Recurrent Neural NetworksPublished by Association for Computational Linguistics (ACL) ,2015
- Large-Scale Video Classification with Convolutional Neural NetworksPublished by Institute of Electrical and Electronics Engineers (IEEE) ,2014
- Meteor Universal: Language Specific Translation Evaluation for Any Target LanguagePublished by Association for Computational Linguistics (ACL) ,2014
- Action Recognition with Improved TrajectoriesPublished by Institute of Electrical and Electronics Engineers (IEEE) ,2013
- Image Classification with the Fisher Vector: Theory and PracticeInternational Journal of Computer Vision, 2013
- Video Google: a text retrieval approach to object matching in videosPublished by Institute of Electrical and Electronics Engineers (IEEE) ,2003
- Long Short-Term MemoryNeural Computation, 1997