Fetching the paper…
Reading the bibliography…
Not all video frames are equally informative for recognizing an action.
Action snippets: How many frames does human action recognition require?
Konrad Schindler and Luc Van Gool · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Hmdb: A large video database for human motion recognition
Hilde Kuehne, Hueihan Jhuang, E. Garrote, T. Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu · 2012
Earlier work this paper cites.
A dataset of 101 human action classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and M Shah · 2012
Earlier work this paper cites.
C3d: Generic features for video analysis
Tran Du, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
H. Kuehne, A. B. Arslan, and T. Serre · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
What do 15,000 object categories tell us about classifying and localizing actions?
Mihir Jain, Jan C Van Gemert, and Cees GM Snoek · 2015
Earlier work this paper cites.
Differential recurrent neural networks for action recognition
Vivek Veeriah, Naifan Zhuang, and Guo-Jun Qi · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Rank pooling for action recognition
Basura Fernando, Efstratios Gavves, José Oramas, Amir Ghodrati, and Tinne Tuytelaars · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
A multi-stream bi-directional recurrent neural network for fine-grained action detection
Bharat Singh, Tim K Marks, Michael Jones, Oncel Tuzel, and Ming Shao · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Two stream lstm: A deep fusion framework for human action recognition
Harshala Gammulle, Simon Denman, Sridha Sridharan, and Clinton Fookes · 2017
Cited alongside, same era.
Actionvlad: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell · 2017
Cited alongside, same era.
The reversible residual network: Backpropagation without storing activations
Aidan N Gomez, Mengye Ren, Raquel Urtasun, and Roger B Grosse · 2017
Cited alongside, same era.
The” something something” video database for learning and evaluating visual common sense
Videograph: Recognizing minutes-long human activities in videos
Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders · 2019
Later among the works it cites.
Stm: Spatiotemporal and motion encoding for action recognition
Boyuan Jiang, MengMeng Wang, Weihao Gan, Wei Wu, and Junjie Yan · 2019
Later among the works it cites.
Scsampler: Sampling salient clips from video for efficient action recognition
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2019
Later among the works it cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Later among the works it cites.
E2-train: Training state-of-the-art cnns with over 80% energy savings
Yue Wang, Ziyu Jiang, Xiaohan Chen, Pengfei Xu, Yang Zhao, Yingyan Lin, and Zhangyang Wang · 2019
Later among the works it cites.
Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Action recognition in video sequences using deep bi-directional lstm with cnn features
Amin Ullah, Jamil Ahmad, Khan Muhammad, Muhammad Sajjad, and Sung Wook Baik · 2017
Cited alongside, same era.
Action recognition with dynamic image networks
H. Bilen, B. Fernando, E. Gavves, and A. Vedaldi · 2018
Cited alongside, same era.
Reversible architectures for arbitrarily deep residual neural networks
Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert, and Elliot Holtham · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
Deep adaptive temporal pooling for activity recognition
Sibo Song, Ngai-Man Cheung, Vijay Chandrasekhar, and Bappaditya Mandal · 2018
Cited alongside, same era.
Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, and Shilei Wen · 2019
Later among the works it cites.
Liteeval: A coarse-to-fine framework for resource efficient video recognition
Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang, and Larry S Davis · 2019
Later among the works it cites.
Adaframe: Adaptive frame selection for fast video recognition
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S. Davis · 2019
Later among the works it cites.
Action recognition with spatial-temporal discriminative filter banks
Brais Martinez, Davide Modolo, Yuanjun Xiong, and Joseph Tighe · 2019
Later among the works it cites.
Resprop: Reuse sparsified backpropagation
Negar Goli and Tor M Aamodt · 2020
Later among the works it cites.
Sideways: Depth-parallel training of video models
Mateusz Malinowski, Grzegorz Swirszcz, Joao Carreira, and Viorica Patraucean · 2020
Later among the works it cites.
Randomized automatic differentiation
Deniz Oktay, Nick McGreivy, Joshua Aduol, Alex Beatson, and Ryan P Adams · 2020
Later among the works it cites.
Best frame selection in a short video
Jian Ren, Xiaohui Shen, Zhe Lin, and Radomir Mech · 2020
Later among the works it cites.
Temporal aggregate representations for long-range video understanding
Fadime Sener, Dipika Singhania, and Angela Yao · 2020
Later among the works it cites.
Gate-shift networks for video action recognition
Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz · 2020
Later among the works it cites.
Dithered backprop: A sparse and quantized backpropagation algorithm for more efficient deep neural network training
Simon Wiedemann, Temesgen Mehari, Kevin Kepp, and Wojciech Samek · 2020
Later among the works it cites.
A multigrid method for efficiently training video models
Chao-Yuan Wu, Ross Girshick, Kaiming He, Christoph Feichtenhofer, and Philipp Krahenbuhl · 2020
Later among the works it cites.
Motionsqueeze: Neural motion feature learning for video understanding
Heeseung Kwon, Manjin Kim, Suha Kwak, and Minsu Cho · 2020
Later among the works it cites.