Fetching the paper…
Reading the bibliography…
How can we collect and use a video dataset to further improve spatiotemporal 3D Convolutional Neural Networks (3D CNNs)? In order to positively answer this open question in video recognition, we have conducted an exploration study using a couple of large-scale video datasets and 3D CNNs.
Space-time Interest Points
I. Laptev and T. Lindeberg · 2003
Earlier work this paper cites.
Actions as space-time shapes
M. Blank, L. Gorelick, E. Shechtman, M. Irani, and R. Basri · 2005
Earlier work this paper cites.
On Space-Time Interest Points
I. Laptev · 2005
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Actions in Context
M. Marszałek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Action Recognition by Dense Trajectories
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2011
Earlier work this paper cites.
UCF101: A Dataset of 101 Human Action Classes From Videos in The Wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
3D Convolutional Neural Networks for Human Action Recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Earlier work this paper cites.
Action Recognition with Improved Trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
DeCAF:A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2014
Earlier work this paper cites.
Large-scale Video Classification with Convolutional Neural Networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Cited alongside, same era.
Two-Stream Convolutional Networks for Action Recognition in Videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Long-term Recurrent Convolutional Networks for Visual Recognition and Description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles · 2015
Cited alongside, same era.
Learning Spatiotemporal Features with 3D Convolutional Networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
The Kinetics Human Action Video Dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, Suleyman M., and A. Zisserman · 2017
Later among the works it cites.
Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
Z. Qiu, T. Yao, and T. Mei · 2017
Later among the works it cites.
ConvNet Architecture Search for Spatiotemporal Feature Learning
D. Tran, J. Ray, Z. Shou, S.-F. Chang, and M. Paluri · 2017
Later among the works it cites.
Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?
K. Hara, H. Kataoka, and Y. Satoh · 2018
Later among the works it cites.
Exploring the Limits of Weakly Supervised Pretraining
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Wang, Y. Xiong, Z. Wang, and Y. Qiao · 2015
Cited alongside, same era.
Youtube-8m: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Cited alongside, same era.
Convolutional Two-Stream Network Fusion for Video Action Recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Cited alongside, same era.
What makes ImageNet good for transfer learning?
M. Huh, P. Agrawal, and A. Efros · 2016
Cited alongside, same era.
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. V. Gool · 2016
Cited alongside, same era.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
A. Carreira and A. Zisserman · 2017
Cited alongside, same era.
A Closer Look at Spatiotemporal Convolutions for Action Recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Later among the works it cites.
STAIR Actions: A Video Dataset of Everyday Home Actions
Y. Yoshikawa, J. Lin, and A. Takeuchi · 2018
Later among the works it cites.
A Short Note on the Kinetics-700 Human Action Dataset
J. Carreira, E. Noland, C. Hillier, and A. Zisserman · 2019
Later among the works it cites.
Holistic Large Scale Video Understanding
A. Diba, M. Fayyaz, V. Sharma, M. Paluri, J. Gall, R. Stiefelhagen, and L. V. Gool · 2019
Later among the works it cites.
Large-scale weakly-supervised pre-training for video action recognition
D. Ghadiyaram, M. Feiszli, D. Tran, X. Yan, H. Wang, and D. Mahajan · 2019
Later among the works it cites.
Do Better ImageNet Models Transfer Better?
S. Kornblith, J. Shlens, and Q. V. Le · 2019
Later among the works it cites.