Fetching the paper…
Reading the bibliography…
Video Captioning and Summarization have become very popular in the recent years due to advancements in Sequence Modelling, with the resurgence of Long-Short Term Memory networks (LSTMs) and introduction of Gated Recurrent Units (GRUs).
Performance of optical flow techniques
John L Barron, David J Fleet, and Steven S Beauchemin · 1994
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Two-frame motion estimation based on polynomial expansion
Gunnar Farnebäck · 2003
Earlier work this paper cites.
High accuracy optical flow estimation based on a theory for warping
Thomas Brox, Andrés Bruhn, Nils Papenberg, and Joachim Weickert · 2004
Earlier work this paper cites.
On space-time interest points
Ivan Laptev · 2005
Earlier work this paper cites.
A 3-dimensional sift descriptor and its application to action recognition
Paul Scovanner, Saad Ali, and Mubarak Shah · 2007
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
Alexander Klaser, Marcin Marszałek, and Cordelia Schmid · 2008
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Dense trajectories and motion boundary descriptors for action recognition
Heng Wang, Alexander Kläser, Cordelia Schmid, and Cheng-Lin Liu · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Grounded language learning from video described with sentences
Haonan Yu and Jeffrey Mark Siskind · 2013
Earlier work this paper cites.
Tv-l1 optical flow estimation
Javier Sánchez Pérez, Enric Meinhardt, and Gabriele Facciolo · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko. Translating videos to natural language using deep recurrent neural networks. In NAACL HLT
2015
Cited alongside, same era.
Sequence to sequence – video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Cited alongside, same era.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Joseph Pal, Hugo Larochelle, and Aaron C. Courville · 2015
Video captioning with transferred semantic attributes
Yingwei Pan, Ting Yao, Houqiang Li, and Tao Mei · 2017
Later among the works it cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Later among the works it cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
The kinetics human action video dataset
Will Kay, João Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Later among the works it cites.
Adaptive pooling in multi-instance learning for web video annotation
Dong Liu, Yizhou Zhou, Xiaoyan Sun, Zheng-Jun Zha, and Wenjun Zeng · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Video paragraph captioning using hierarchical recurrent neural networks
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Cited alongside, same era.
Activitynet: A large-scale video benchmark for human activity understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles · 2015
Cited alongside, same era.
Deep multiple instance learning for image classification and auto-annotation
Jiajun Wu, Yinan Yu, Chang Huang, and Kai Yu · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Cited alongside, same era.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo · 2016
Cited alongside, same era.
Multimodal keyless attention fusion for video classification
Xiang Long, Chuang Gan, Gerard de Melo, Xiao Liu, Yandong Li, Fu Li, and Shilei Wen · 2018
Later among the works it cites.
Video captioning with multi-faceted attention
Xiang Long, Chuang Gan, and Gerard de Melo · 2018
Later among the works it cites.
End-to-end dense video captioning with masked transformer
Luowei Zhou, Yingbo Zhou, Jason J. Corso, Richard Socher, and Caiming Xiong · 2018
Later among the works it cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Closest in time.