Fetching the paper…
Reading the bibliography…
The work in this paper is driven by the question how to exploit the temporal cues available in videos for their accurate classification, and for human action recognition in particular? Thus far, the vision community has focused on spatio-temporal approaches with fixed temporal convolution kernel depths.
Human detection using oriented histograms of flow and appearance
N. Dalal, B. Triggs, and C. Schmid · 2006
Earlier work this paper cites.
A 3-dimensional sift descriptor and its application to action recognition
P. Scovanner, S. Ali, and M. Shah · 2007
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
A. Klaser, M. Marszałek, and C. Schmid · 2008
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
An efficient dense and scale-invariant spatio-temporal interest point detector
G. Willems, T. Tuytelaars, and L. Van Gool · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Hmdb51: A large video database for human motion recognition
H. Kuehne, H. Jhuang, R. Stiefelhagen, and T. Serre · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Multi-view super vector for action recognition
Z. Cai, L. Wang, X. Peng, and Y. Qiao · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Modeling video evolution for action recognition
B. Fernando, E. Gavves, J. M. Oramas, A. Ghodrati, and T. Tuytelaars · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Network in network
M. Lin, Q. Chen, and S. Yan · 2015
Spatiotemporal residual networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. Wildes · 2016
Later among the works it cites.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Later among the works it cites.
Cross modal distillation for supervision transfer
S. Gupta, J. Hoffman, and J. Malik · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Infrared colorization using deep convolutional neural networks
M. Limmer and H. P. Lensch · 2016
Later among the works it cites.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human action recognition using factorized spatio-temporal convolutional networks
L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Cited alongside, same era.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Cited alongside, same era.
Youtube-8m: A large-scale video classification benchmark
S. Abu-El-Haija, N. Kothari, J. Lee, P. Natsev, G. Toderici, B. Varadarajan, and S. Vijayanarasimhan · 2016
Cited alongside, same era.
Efficient two-stream motion and appearance 3d cnns for video classification
A. Diba, A. M. Pazandeh, and L. Van Gool · 2016
Cited alongside, same era.
R. Arandjelović and A. Zisserman · 2017
Closest in time.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Closest in time.
Deep temporal linear encoding networks
A. Diba, V. Sharma, and L. Van Gool · 2017
Closest in time.
Learning spatio-temporal features with 3d residual networks for action recognition
K. Hara, H. Kataoka, and Y. Satoh · 2017
Closest in time.
Densely connected convolutional networks
G. Huang, Z. Liu, K. Q. Weinberger, and L. van der Maaten · 2017
Closest in time.
Convnet architecture search for spatiotemporal feature learning
D. Tran, J. Ray, Z. Shou, S.-F. Chang, and M. Paluri · 2017
Closest in time.