Fetching the paper…
Reading the bibliography…
The purpose of this study is to determine whether current video datasets have sufficient data for training very deep convolutional neural networks (CNNs) with spatio-temporal three-dimensional (3D) kernels.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human action classes from videos in the wild
K. Soomro, A. Roshan Zamir, and M. Shah · 2012
Earlier work this paper cites.
3D convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Dense trajectories and motion boundary descriptors for action recognition
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
ActivityNet: A large-scale video benchmark for human activity understanding
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3D convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Cited alongside, same era.
Towards good practices for very deep two-stream convnets
L. Wang, Y. Xiong, Z. Wang, and Y. Qiao · 2015
Cited alongside, same era.
YouTube-8M: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Cited alongside, same era.
Spatiotemporal residual networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. Wildes · 2016
Cited alongside, same era.
S. Zagoruyko and N. Komodakis · 2016
Later among the works it cites.
Quo vadis, action recognition? A new model and the Kinetics dataset
J. Carreira and A. Zisserman · 2017
Closest in time.
Quo vadis, action recognition? A new model and the Kinetics dataset
J. Carreira and A. Zisserman · 2017
Closest in time.
Spatiotemporal multiplier networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. P. Wildes · 2017
Closest in time.
Learning spatio-temporal features with 3D residual networks for action recognition
K. Hara, H. Kataoka, and Y. Satoh · 2017
Closest in time.
Densely connected convolutional networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Long-term temporal convolutions for action recognition
G. Varol, I. Laptev, and C. Schmid · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Cited alongside, same era.
G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger · 2017
Closest in time.
The Kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Closest in time.
Learning spatio-temporal representation with pseudo-3d residual networks
Z. Qiu, T. Yao, and T. Mei · 2017
Closest in time.
Convnet architecture search for spatiotemporal feature learning
D. Tran, J. Ray, Z. Shou, S. Chang, and M. Paluri · 2017
Closest in time.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Closest in time.