Fetching the paper…
Reading the bibliography…
Typical human actions last several seconds and exhibit characteristic spatio-temporal structure.
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. Howard, W. Hubbard, and L. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural Computation , vol. 1, no. 4, pp. 541–551, 1989
1989
Earlier work this paper cites.
B. Tversky, J. Morrison, and J. Zacks, “On bodies and events,” in The Imitative Mind , A. Meltzoff and W. Prinz, Eds. Cambridge University Press, 2002
2002
Earlier work this paper cites.
G. Farnebäck, “Two-frame motion estimation based on polynomial expansion,” in SCIA , 2003
2003
Earlier work this paper cites.
C. Schüldt, I. Laptev, and B. Caputo, “Recognizing human actions: a local SVM approach,” in ICPR , 2004
2004
Earlier work this paper cites.
G. Csurka, C. Dance, L. Fan, J. Willamowski, and C. Bray, “Visual categorization with bags of keypoints,” in ECCVW , 2004
2004
Earlier work this paper cites.
T. Brox, A. Bruhn, N. Papenberg, and J. Weickert, “High accuracy optical flow estimation based on a theory for warping,” in ECCV , 2004
2004
Earlier work this paper cites.
I. Laptev, M. Marszałek, C. Schmid, and B. Rozenfeld, “Learning realistic human actions from movies,” in CVPR , 2008
2008
Earlier work this paper cites.
J. C. Niebles, H. Wang, and L. Fei-Fei, “Unsupervised learning of human action categories using spatial-temporal words,” IJCV , vol. 79, no. 3, pp. 299–318, 2008
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in CVPR , 2009
2009
Earlier work this paper cites.
S. Ji, W. Xu, M. Yang, and K. Yu, “3D convolutional neural networks for human action recognition,” in ICML , 2010
2010
Earlier work this paper cites.
G. W. Taylor, R. Fergus, Y. LeCun, and C. Bregler, “Convolutional learning of spatio-temporal features,” in ECCV , 2010
2010
Earlier work this paper cites.
F. Perronnin, J. Sánchez, and T. Mensink, “Improving the Fisher kernel for large-scale image classification,” in ECCV , 2010
2010
Earlier work this paper cites.
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “HMDB: a large video database for human motion recognition,” in ICCV , 2011
2011
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in NIPS , 2012
2012
Cited alongside, same era.
K. Soomro, A. Roshan Zamir, and M. Shah, “UCF101: A dataset of 101 human actions classes from videos in the wild,” in CRCV-TR-12-01 , 2012
2012
Cited alongside, same era.
H. Wang and C. Schmid, “Action recognition with improved trajectories,” in ICCV , 2013
2013
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in NIPS , 2014
2014
Cited alongside, same era.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3D convolutional networks,” in ICCV , 2015
2015
Later among the works it cites.
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR , 2015
2015
Later among the works it cites.
B. Fernando, E. Gavves, J. Oramas, A. Ghodrati, and T. Tuytelaars, “Modeling video evolution for action recognition,” in CVPR , 2015
2015
Later among the works it cites.
L. Wang, Y. Qiao, and X. Tang, “Action recognition with trajectory-pooled deep-convolutional descriptors,” in CVPR , 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learning deep features for scene recognition using places database,” in NIPS , 2014
2014
Cited alongside, same era.
R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in CVPR , 2014
2014
Cited alongside, same era.
Y. Taigman, M. Yang, M. Ranzato, and L. Wolf, “DeepFace: Closing the gap to human-level performance in face verification,” in CVPR , 2014
2014
Cited alongside, same era.
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” in CVPR , 2014
2014
Cited alongside, same era.
V. Kantorov and I. Laptev, “Efficient feature extraction, encoding, and classification for action recognition,” in CVPR , 2014
2014
Cited alongside, same era.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in ECCV , 2014
2014
Cited alongside, same era.
http://www.di.ens.fr/willow/research/ltc/
Cited in the paper.
2015
Later among the works it cites.
Z.-Z. Lan, M. Lin, X. Li, A. G. Hauptmann, and B. Raj, “Beyond Gaussian pyramid: Multi-skip feature stacking for action recognition.” in CVPR , 2015
2015
Later among the works it cites.
J. Y. Ng, M. J. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici, “Beyond short snippets: Deep networks for video classification,” in CVPR , 2015
2015
Later among the works it cites.
H. Bilen, B. Fernando, E. Gavves, A. Vedaldi, and S. Gould, “Dynamic image networks for action recognition,” in CVPR , 2016
2016
Closest in time.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” in CVPR , 2016
2016
Closest in time.
B. Zhang, L. Wang, Z. Wang, Y. Qiao, and H. Wang, “Real-time action recognition with enhanced motion vector CNNs,” in CVPR , 2016
2016
Closest in time.
X. Wang, A. Farhadi, and A. Gupta, “Actions ~ transformations,” in CVPR , 2016
2016
Closest in time.