Fetching the paper…
Reading the bibliography…
The paucity of videos in current action classification datasets (UCF-101 and HMDB-51) has made it difficult to identify good video architectures, as most methods obtain similar performance on existing small-scale benchmarks.
The recognition of human movement using temporal templates
A. F. Bobick and J. W. Davis · 2001
Earlier work this paper cites.
A duality based approach for realtime TV-L1 optical flow
C. Zach, T. Pock, and H. Bischof · 2007
Earlier work this paper cites.
Action recognition by learning mid-level motion features
A. Fathi and G. Mori · 2008
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Unsupervised learning of human action categories using spatial-temporal words
J. C. Niebles, H. Wang, and L. Fei-Fei · 2008
Earlier work this paper cites.
Human focused action localization in video
A. Kläser, M. Marszalek, C. Schmid, and A. Zisserman · 2010
Earlier work this paper cites.
Convolutional learning of spatio-temporal features
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler · 2010
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Learning and transferring mid-level image representations using convolutional neural networks
M. Oquab, L. Bottou, I. Laptev, and J. Sivic · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Modeling video evolution for action recognition
B. Fernando, E. Gavves, J. M. Oramas, A. Ghodrati, and T. Tuytelaars · 2015
Cited alongside, same era.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al · 2016
Later among the works it cites.
Dynamic image networks for action recognition
H. Bilen, B. Fernando, E. Gavves, A. Vedaldi, and S. Gould · 2016
Later among the works it cites.
T. Cooijmans, N. Ballas, C. Laurent, and A. Courville · 2016
Later among the works it cites.
Spatiotemporal residual networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. P. Wildes · 2016
Later among the works it cites.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Initialization strategies of spatio-temporal convolutional neural networks
E. Mansimov, N. Srivastava, and R. Salakhutdinov · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Cited alongside, same era.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
VideoLSTM convolves, attends and flows for action recognition
Z. Li, E. Gavves, M. Jain, and C. G. Snoek · 2016
Later among the works it cites.
Multi-region two-stream R-CNN for action detection
X. Peng and C. Schmid · 2016
Later among the works it cites.
Deep learning for detecting multiple space-time action tubes in videos
S. Saha, G. Singh, M. Sapienza, P. H. Torr, and F. Cuzzolin · 2016
Later among the works it cites.
Temporal segment networks: towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Later among the works it cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Closest in time.
Long-term temporal convolutions for action recognition
G. Varol, I. Laptev, and C. Schmid · 2017
Closest in time.