Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Observing human-object interactions: Using spatial and functional compatibility for recognition
Abhinav Gupta, Aniruddha Kembhavi, and Larry S Davis · 2009
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
Bangpeng Yao and Li Fei-Fei · 2010
Earlier work this paper cites.
Semantic segmentation with second-order pooling
Joao Carreira, Rui Caseiro, Jorge Batista, and Cristian Sminchisescu · 2012
Earlier work this paper cites.
Describing the scene as a whole: Joint object detection, scene classification and semantic segmentation
Jian Yao, Sanja Fidler, and Raquel Urtasun · 2012
Earlier work this paper cites.
Temporal localization of actions with actoms
Adrien Gaidon, Zaid Harchaoui, and Cordelia Schmid · 2013
Earlier work this paper cites.
Action and event recognition with fisher vectors on a compact feature set
Dan Oneata, Jakob Verbeek, and Cordelia Schmid · 2013
Earlier work this paper cites.
Combining the right features for complex event recognition
Kevin Tang, Bangpeng Yao, Li Fei-Fei, and Daphne Koller · 2013
Earlier work this paper cites.
Action localization with tubelets from motion
Mihir Jain, Jan Van Gemert, Hervé Jégou, Patrick Bouthemy, and Cees GM Snoek · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Bag-of-fragments: Selecting and encoding video fragments for event detection and recounting
Pascal Mettes, Jan C van Gemert, Spencer Cappallo, Thomas Mensink, and Cees GM Snoek · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Towards good practices for very deep two-stream convnets
Original
Limin Wang, Yuanjun Xiong, Zhe Wang, and Yu Qiao · 2015
Earlier work this paper cites.
Tensorflow: a system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.