Fetching the paper…
Reading the bibliography…
Deep convolutional networks have achieved great success for image recognition.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” in CVPR , 2016, pp. 1933–1941
1941
Earlier work this paper cites.
B. K. Horn and B. G. Schunck, “Determining optical flow,” Artificial intelligence , vol. 17, no. 1-3, pp. 185–203, 1981
1981
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
W. Zhu, J. Hu, G. Sun, X. Cao, and Y. Qiao, “A key volume mining deep framework for action recognition,” in CVPR , 2016, pp. 1991–1999
1999
Earlier work this paper cites.
G. Csurka, C. Dance, L. Fan, J. Willamowski, and C. Bray, “Visual categorization with bags of keypoints,” in ECCV Workshop on statistical learning in computer vision , 2004, pp. 1–22
2004
Earlier work this paper cites.
D. A. Forsyth, O. Arikan, L. Ikemoto, J. F. O’Brien, and D. Ramanan, “Computational studies of human motion: Part 1, tracking and motion synthesis,” Foundations and Trends in Computer Graphics and Vision , vol. 1, no. 2/3, 2005
2005
Earlier work this paper cites.
I. Laptev, “On space-time interest points,” International Journal of Computer Vision , vol. 64, no. 2-3, pp. 107–123, 2005
2005
Earlier work this paper cites.
P. Dollár, V. Rabaud, G. Cottrell, and S. Belongie, “Behavior recognition via sparse spatio-temporal features,” in IEEE International Workshop on PETS , 2005
2005
Earlier work this paper cites.
C. Zach, T. Pock, and H. Bischof, “A duality based approach for realtime tv- L 1 L^{1} optical flow,” in 29th DAGM Symposium on Pattern Recognition , 2007, pp. 214–223
2007
Earlier work this paper cites.
P. K. Turaga, R. Chellappa, V. S. Subrahmanian, and O. Udrea, “Machine recognition of human activities: A survey,” IEEE Trans. Circuits Syst. Video Techn. , vol. 18, no. 11, pp. 1473–1488, 2008
2008
Earlier work this paper cites.
G. Willems, T. Tuytelaars, and L. J. V. Gool, “An efficient dense and scale-invariant spatio-temporal interest point detector,” in ECCV , 2008, pp. 650–663
2008
Earlier work this paper cites.
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld, “Learning realistic human actions from movies,” in CVPR , 2008, pp. 1–8
2008
Earlier work this paper cites.
A. Kläser, M. Marszalek, and C. Schmid, “A spatio-temporal descriptor based on 3D-gradients,” in BMVC , 2008, pp. 1–12
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li, “ImageNet: A large-scale hierarchical image database,” in CVPR , 2009, pp. 248–255
2009
Earlier work this paper cites.
L. D. Bourdev and J. Malik, “Poselets: Body part detectors trained using 3d human pose annotations,” in ICCV , 2009, pp. 1365–1372
2009
Earlier work this paper cites.
J. C. Niebles, C.-W. Chen, and F.-F. Li, “Modeling temporal structure of decomposable motion segments for activity classification,” in ECCV , 2010, pp. 392–405
2010
Earlier work this paper cites.
P. F. Felzenszwalb, R. B. Girshick, D. A. McAllester, and D. Ramanan, “Object detection with discriminatively trained part-based models,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 32, no. 9, pp. 1627–1645, 2010
2010
Earlier work this paper cites.
H. Kuehne, H. Jhuang, E. Garrote, T. A. Poggio, and T. Serre, “HMDB: A large video database for human motion recognition,” in ICCV , 2011, pp. 2556–2563
2011
Earlier work this paper cites.
J. K. Aggarwal and M. S. Ryoo, “Human activity analysis: A review,” ACM Comput. Surv. , vol. 43, no. 3, p. 16, 2011
2011
Earlier work this paper cites.
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu, “Action recognition by dense trajectories,” in CVPR , 2011, pp. 3169–3176
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in NIPS , 2012, pp. 1106–1114
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
H. Jégou, F. Perronnin, M. Douze, J. Sánchez, P. Pérez, and C. Schmid, “Aggregating local image descriptors into compact codes,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 34, no. 9, pp. 1704–1716, 2012
2012
Earlier work this paper cites.
M. Raptis, I. Kokkinos, and S. Soatto, “Discovering discriminative action parts from mid-level video representations,” in CVPR , 2012, pp. 1242–1249
2012
Earlier work this paper cites.
S. Sadanand and J. J. Corso, “Action bank: A high-level representation of activity in video,” in CVPR , 2012, pp. 1234–1241
2012
Earlier work this paper cites.
H. Wang and C. Schmid, “Action recognition with improved trajectories,” in ICCV , 2013, pp. 3551–3558
2013
Cited alongside, same era.
L. Wang, Y. Qiao, and X. Tang, “Motionlets: Mid-level 3D parts for human motion recognition,” in CVPR , 2013, pp. 2674–2681
2013
Cited alongside, same era.
A. Gaidon, Z. Harchaoui, and C. Schmid, “Temporal localization of actions with actoms,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 35, no. 11, pp. 2782–2795, 2013
2013
Cited alongside, same era.
J. Sánchez, F. Perronnin, T. Mensink, and J. J. Verbeek, “Image classification with the fisher vector: Theory and practice,” International Journal of Computer Vision , vol. 105, no. 3, pp. 222–245, 2013
2013
Cited alongside, same era.
A. Jain, A. Gupta, M. Rodriguez, and L. S. Davis, “Representing videos using mid-level discriminative patches,” in CVPR , 2013, pp. 2571–2578
B. Fernando, E. Gavves, J. O. M., A. Ghodrati, and T. Tuytelaars, “Modeling video evolution for action recognition,” in CVPR , 2015, pp. 5378–5387
2015
Later among the works it cites.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR , 2015, pp. 2625–2634
2015
Later among the works it cites.
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in CVPR , 2015, pp. 961–970
2015
Later among the works it cites.
L. Sun, K. Jia, D. Yeung, and B. E. Shi, “Human action recognition using factorized spatio-temporal convolutional networks,” in ICCV , 2015, pp. 4597–4605
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
W. Zhang, M. Zhu, and K. G. Derpanis, “From actemes to action: A strongly-supervised representation for detailed action understanding,” in ICCV , 2013, pp. 2248–2255
2013
Cited alongside, same era.
J. Zhu, B. Wang, X. Yang, W. Zhang, and Z. Tu, “Action recognition with actons,” in ICCV , 2013, pp. 3559–3566
2013
Cited alongside, same era.
S. Ji, W. Xu, M. Yang, and K. Yu, “3D convolutional neural networks for human action recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 35, no. 1, pp. 221–231, 2013
2013
Cited alongside, same era.
Y.-G. Jiang, J. Liu, A. Roshan Zamir, I. Laptev, M. Piccardi, M. Shah, and R. Sukthankar, “THUMOS challenge: Action recognition with a large number of classes,” 2013
2013
Cited alongside, same era.
H. Wang and C. Schmid, “LEAR-INRIA submission for the thumos workshop,” in ICCV Workshop on THUMOS Challenge , 2013, pp. 1–3
2013
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in NIPS , 2014, pp. 568–576
2014
Cited alongside, same era.
B. Zhou, À. Lapedriza, J. Xiao, A. Torralba, and A. Oliva, “Learning deep features for scene recognition using places database,” in NIPS , 2014, pp. 487–495
2014
Cited alongside, same era.
2015
Later among the works it cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015, pp. 448–456
2015
Later among the works it cites.
2015
Later among the works it cites.
M. Jain, J. C. van Gemert, and C. G. Snoek, “What do 15,000 object categories tell us about classifying and localizing actions?” in CVPR , 2015, pp. 46–55
2015
Later among the works it cites.
B. Ni, P. Moulin, X. Yang, and S. Yan, “Motion part regularization: Improving action recognition via trajectory group selection,” in CVPR , 2015, pp. 3698–3706
2015
Later among the works it cites.
L. Shen, Z. Lin, and Q. Huang, “Relay backpropagation for effective learning of deep convolutional neural networks,” in ECCV , 2016, pp. 467–482
2016
Later among the works it cites.
2016
Later among the works it cites.
B. Zhang, L. Wang, Z. Wang, Y. Qiao, and H. Wang, “Real-time action recognition with enhanced motion vector CNNs,” in CVPR , 2016, pp. 2718–2726
2016
Later among the works it cites.
2016
Later among the works it cites.
H. Idrees, A. R. Zamir, Y.-G. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah, “The THUMOS challenge on action recognition for videos “in the wild”,” Computer Vision and Image Understanding , pp. –, 2016
2016
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Later among the works it cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in CVPR , 2016, pp. 2818–2826
2016
Later among the works it cites.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good practices for deep action recognition,” in ECCV , 2016, pp. 20–36
2016
Later among the works it cites.
L. Wang, Y. Qiao, and X. Tang, “MoFAP: A multi-level representation for action recognition,” International Journal of Computer Vision , vol. 119, no. 3, pp. 254–271, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
C. Feichtenhofer, A. Pinz, and R. P. Wildes, “Spatiotemporal residual networks for video action recognition,” in NIPS , 2016, pp. 3468–3476
2016
Later among the works it cites.
Y. Zhu and S. Newsam, “Depth2action: Exploring embedded depth for large-scale action recognition,” in ECCV . Springer, 2016, pp. 668–684
2016
Later among the works it cites.
X. Peng, L. Wang, X. Wang, and Y. Qiao, “Bag of visual words and fusion methods for action recognition: Comprehensive study and good practice,” Computer Vision and Image Understanding , vol. 150, pp. 109–125, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
L. Wang, S. Guo, W. Huang, Y. Xiong, and Y. Qiao, “Knowledge guided disambiguation for large-scale scene classification with multi-resolution cnns,” IEEE Trans. Image Processing , vol. 26, no. 4, pp. 2055–2068, 2017
2017
Closest in time.