Fetching the paper…
Reading the bibliography…
Typical video classification methods often divide a video into short clips, do inference on each clip independently, then aggregate the clip-level predictions to generate the video-level results.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
On space-time interest points
I. Laptev · 2005
Earlier work this paper cites.
Human detection using oriented histograms of flow and appearance
N. Dalal, B. Triggs, and C. Schmid · 2006
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
A. Klaser, M. Marszałek, and C. Schmid · 2008
Earlier work this paper cites.
Sequential deep learning for human action recognition
M. Baccouche, F. Mamalet, C. Wolf, C. Garcia, and A. Baskurt · 2011
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Action recognition by dense trajectories
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Unsupervised learning of video representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Cited alongside, same era.
Human action recognition using factorized spatio-temporal convolutional networks
L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3D convolutional networks
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Later among the works it cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Z. Qiu, T. Yao, and T. Mei · 2017
Later among the works it cites.
Appearance-and-relation networks for video classification
L. Wang, W. Li, W. Li, and L. Van Gool · 2017
Later among the works it cites.
Action search: Spotting actions in videos and its application to temporal action localization
H. Alwassel, F. Caba Heilbron, and B. Ghanem · 2018
Later among the works it cites.
Multi-fiber networks for video recognition
Y. Chen, Y. Kalantidis, J. Li, S. Yan, and J. Feng · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Cited alongside, same era.
Netvlad: Cnn architecture for weakly supervised place recognition
R. Arandjelovic, P. Gronat, A. Torii, T. Pajdla, and J. Sivic · 2016
Cited alongside, same era.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Cited alongside, same era.
A. Diba, M. Fayyaz, V. Sharma, M. M. Arzani, R. Yousefzadeh, J. Gall, and L. Van Gool · 2018
Later among the works it cites.
Videolstm convolves, attends and flows for action recognition
Z. Li, K. Gavrilyuk, E. Gavves, M. Jain, and C. G. Snoek · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Later among the works it cites.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Later among the works it cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy · 2018
Later among the works it cites.
Mict: Mixed 3d/2d convolutional tube for human action recognition
Y. Zhou, X. Sun, Z.-J. Zha, and W. Zeng · 2018
Later among the works it cites.
Eco: Efficient convolutional network for online video understanding
M. Zolfaghari, K. Singh, and T. Brox · 2018
Later among the works it cites.
Efficient video classification using fewer frames
S. Bhardwaj, M. Srinivasan, and M. M. Khapra · 2019
Closest in time.
Scsampler: Sampling salient clips from video for efficient action recognition
B. Korbar, D. Tran, and L. Torresani · 2019
Closest in time.
Adaframe: Adaptive frame selection for fast video recognition
Z. Wu, C. Xiong, C.-Y. Ma, R. Socher, and L. S. Davis · 2019
Closest in time.