Fetching the paper…
Reading the bibliography…
We propose an approach to learn spatio-temporal features in videos from intermediate visual representations we call "percepts" using Gated-Recurrent-Unit Recurrent Networks (GRUs).Our method relies on percepts that are extracted from all level of a deep convolutional network trained on the large ImageNet dataset.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, Kishore, Roukos, Salim, Ward, Todd, and Zhu, Wei-Jing · 2002
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, James, Breuleux, Olivier, Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Desjardins, Guillaume, Turian, Joseph, Warde-Farley, David, and Bengio, Yoshua · 2010
Earlier work this paper cites.
Large displacement optical flow: descriptor matching in variational motion estimation
Brox, T. and Malik, J · 2011
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, David L and Dolan, William B · 2011
Earlier work this paper cites.
Action recognition by dense trajectories
Wang, H., Kläser, A., Schmid, C., and Liu, C · 2011
Earlier work this paper cites.
Theano: new features and speed improvements
Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian J., Bergeron, Arnaud, Bouchard, Nicolas, and Bengio, Yoshua · 2012
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, James and Bengio, Yoshua · 2012
Earlier work this paper cites.
Action bank: A high-level representation of activity in video
Sadanand, S. and Corso, J · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, Khurram, Zamir, Amir Roshan, and Shah, Mubarak · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Zeiler, Matthew D · 2012
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, Junyoung, Gulcehre, Caglar, Cho, KyungHyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Meteor universal: Language specific translation evaluation for any target language
Denkowski, Michael and Lavie, Alon · 2014
Cited alongside, same era.
Integrating language and vision to generate natural language descriptions of videos in the wild
Thomason, Jesse, Venugopalan, Subhashini, Guadarrama, Sergio, Saenko, Kate, and Mooney, Raymond · 2014
Later among the works it cites.
C3d: generic features for video analysis
Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M · 2014
Later among the works it cites.
CIDEr: Consensus-based image description evaluation
Vedantam, Ramakrishna, Zitnick, C Lawrence, and Parikh, Devi · 2014
Later among the works it cites.
Microsoft coco captions: Data collection and evaluation server
Chen, Xinlei, Fang, Hao, Lin, Tsung-Yi, Vedantam, Ramakrishna, Gupta, Saurabh, Dollar, Piotr, and Zitnick, C Lawrence · 2015
Closest in time.
Beyond short snippets: Deep networks for video classification
Ng, Joe Yue-Hei, Hausknecht, Matthew, Vijayanarasimhan, Sudheendra, Vinyals, Oriol, Monga, Rajat, and Toderici, George · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., and Darrell, T · 2014
Cited alongside, same era.
Towards end-to-end speech recognition with recurrent neural networks
Graves, A. and Jaitly, N · 2014
Cited alongside, same era.
Thumos challenge: Action recognition with a large number of classes
Jiang, YG, Liu, J, Roshan Zamir, A, Toderici, G, Laptev, I, Shah, M, and Sukthankar, R · 2014
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
Karpathy, Andrej, Toderici, George, Shetty, Sachin, Leung, Tommy, Sukthankar, Rahul, and Fei-Fei, Li · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Beyond gaussian pyramid: Multi-skip feature stacking for action recognition
Lan, Zhenzhong, Lin, Ming, Li, Xuanchong, Hauptmann, Alexander G, and Raj, Bhiksha · 2014
Cited alongside, same era.
Going deeper with convolutions
Szegedy, Christian, Liu, Wei, Jia, Yangqing, Sermanet, Pierre, Reed, Scott, Anguelov, Dragomir, Erhan, Dumitru, Vanhoucke, Vincent, and Rabinovich, Andrew · 2014
Cited alongside, same era.
Closest in time.
Hierarchical recurrent neural encoder for video representation with application to captioning
Pan, Pingbo, Xu, Zhongwen, Yang, Yi, Wu, Fei, and Zhuang, Yueting · 2015
Closest in time.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Shi, Xingjian, Chen, Zhourong, Wang, Hao, Yeung, Dit-Yan, Wong, Wai-Kin, and Woo, Wang-chun · 2015
Closest in time.
Unsupervised learning of video representations using lstms
Srivastava, N., Mansimov, E., and Salakhutdinov, R · 2015
Closest in time.
Translating videos to natural language using deep recurrent neural networks
Venugopalan, Subhashini, Xu, Huijuan, Donahue, Jeff, Rohrbach, Marcus, Mooney, Raymond, and Saenko, Kate · 2015
Closest in time.
Describing videos by exploiting temporal structure
Yao, Li, Torabi, Atousa, Cho, Kyunghyun, Ballas, Nicolas, Pal, Christopher, Larochelle, Hugo, and Courville, Aaron · 2015
Closest in time.
Video paragraph captioning using hierarchical recurrent neural networks
Yu, Haonan, Wang, Jiang, Huang, Zhiheng, Yang, Yi, and Xu, Wei · 2015
Closest in time.