Fetching the paper…
Reading the bibliography…
Different from RGB videos, depth data in RGB-D videos provide key complementary information for tristimulus visual data which potentially could achieve accuracy improvement for action recognition.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 1933–1941
1941
Earlier work this paper cites.
W. Zhu, J. Hu, G. Sun, X. Cao, and Y. Qiao, “A key volume mining deep framework for action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 1991–1999
1999
Earlier work this paper cites.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in Computer Vision and Pattern Recognition, 2005. CVPR 2005. IEEE Computer Society Conference on , vol. 1. IEEE, 2005, pp. 886–893
2005
Earlier work this paper cites.
N. Dalal, B. Triggs, and C. Schmid, “Human detection using oriented histograms of flow and appearance,” in European conference on computer vision . Springer, 2006, pp. 428–441
2006
Earlier work this paper cites.
F. Lv and R. Nevatia, “Recognition and segmentation of 3-d human action using hmm and multi-class adaboost,” in European conference on computer vision . Springer, 2006, pp. 359–372
2006
Earlier work this paper cites.
M. Müller and T. Röder, “Motion templates for automatic classification and retrieval of motion capture data,” in Proceedings of the 2006 ACM SIGGRAPH/Eurographics symposium on Computer animation . Eurographics Association, 2006, pp. 137–146
2006
Earlier work this paper cites.
P. Scovanner, S. Ali, and M. Shah, “A 3-dimensional sift descriptor and its application to action recognition,” in Proceedings of the 15th ACM international conference on Multimedia . ACM, 2007, pp. 357–360
2007
Earlier work this paper cites.
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld, “Learning realistic human actions from movies,” in Computer Vision and Pattern Recognition, 2008. CVPR 2008. IEEE Conference on . IEEE, 2008, pp. 1–8
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
L. Han, X. Wu, W. Liang, G. Hou, and Y. Jia, “Discriminative human action recognition in the learned hierarchical manifold space,” Image and Vision Computing , vol. 28, no. 5, pp. 836–849, 2010
2010
Earlier work this paper cites.
R. Marks, “System and method for providing a real-time three-dimensional interactive environment,” Dec. 6 2011, uS Patent 8,072,470
2011
Earlier work this paper cites.
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu, “Action recognition by dense trajectories,” in Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on . IEEE, 2011, pp. 3169–3176
2011
Earlier work this paper cites.
M. Baccouche, F. Mamalet, C. Wolf, C. Garcia, and A. Baskurt, “Sequential deep learning for human action recognition,” in International Workshop on Human Behavior Understanding . Springer, 2011, pp. 29–39
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Proc. Adv. Neural Inf. Process. Syst. , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
Z. Zhang, “Microsoft kinect sensor and its effect,” IEEE multimedia , vol. 19, no. 2, pp. 4–10, 2012
2012
Earlier work this paper cites.
J. Wang, Z. Liu, Y. Wu, and J. Yuan, “Mining actionlet ensemble for action recognition with depth cameras,” in Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on . IEEE, 2012, pp. 1290–1297
2012
Earlier work this paper cites.
X. Yang, C. Zhang, and Y. Tian, “Recognizing actions using depth motion maps-based histograms of oriented gradients,” in Proceedings of the 20th ACM international conference on Multimedia . ACM, 2012, pp. 1057–1060
2012
Earlier work this paper cites.
R. Li and T. Zickler, “Discriminative virtual views for cross-view action recognition,” in Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on . IEEE, 2012, pp. 2855–2862
2012
Earlier work this paper cites.
J. Yu, K. Weng, G. Liang, and G. Xie, “A vision-based robotic grasping system using deep learning for 3d object recognition and pose estimation,” in Robotics and Biomimetics (ROBIO), 2013 IEEE International Conference on . IEEE, Dec. 2013, pp. 1175–1180
2013
Earlier work this paper cites.
A. Gaidon, Z. Harchaoui, and C. Schmid, “Temporal localization of actions with actoms,” IEEE transactions on pattern analysis and machine intelligence , p. 1, 2013
2013
Earlier work this paper cites.
H. Wang and C. Schmid, “Action recognition with improved trajectories,” in Proceedings of the IEEE international conference on computer vision , 2013, pp. 3551–3558
2013
Earlier work this paper cites.
T. Basha, Y. Moses, and N. Kiryati, “Multi-view scene flow estimation: A view centered variational approach,” International journal of computer vision , vol. 101, no. 1, pp. 6–21, 2013
2013
Earlier work this paper cites.
O. Oreifej and Z. Liu, “Hon4d: Histogram of oriented 4d normals for activity recognition from depth sequences,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2013, pp. 716–723
2013
Earlier work this paper cites.
J. Han, L. Shao, D. Xu, and J. Shotton, “Enhanced computer vision with microsoft kinect sensor: A review,” IEEE transactions on cybernetics , vol. 43, no. 5, pp. 1318–1334, 2013
2013
Earlier work this paper cites.
Z. Zhang, C. Wang, B. Xiao, W. Zhou, S. Liu, and C. Shi, “Cross-view action recognition via a continuous virtual path,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2013, pp. 2690–2697
2013
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in Advances in neural information processing systems , 2014, pp. 568–576
2014
Cited alongside, same era.
H. Rahmani, A. Mahmood, D. Q. Huynh, and A. Mian, “Hopc: Histogram of oriented principal components of 3d pointclouds for action recognition,” in European conference on computer vision . Springer, 2014, pp. 742–757
2014
Cited alongside, same era.
J. Wang, X. Nie, Y. Xia, Y. Wu, and S.-C. Zhu, “Cross-view action modeling, learning and recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2014, pp. 2649–2656
2014
Cited alongside, same era.
A. Gupta, J. Martinez, J. J. Little, and R. J. Woodham, “3d pose from motion for cross-view action recognition via non-linear circulant temporal encoding,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 2601–2608
2014
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Z. Qiu, T. Yao, and T. Mei, “Learning spatio-temporal representation with pseudo-3d residual networks,” in 2017 IEEE International Conference on Computer Vision (ICCV) . IEEE, 2017, pp. 5534–5542
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Wang, Z. Liu, Y. Wu, and J. Yuan, “Learning actionlet ensemble for 3d human action recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 36, no. 5, pp. 914–927, 2014
2014
Cited alongside, same era.
E. Ahmed, M. Jones, and T. K. Marks, “An improved deep learning architecture for person re-identification,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 3908–3916
2015
Cited alongside, same era.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 2625–2634
2015
Cited alongside, same era.
J.-F. Hu, W.-S. Zheng, J. Lai, and J. Zhang, “Jointly learning heterogeneous features for rgb-d activity recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 5344–5352
2015
Cited alongside, same era.
2015
Cited alongside, same era.
L. Wang, Y. Qiao, and X. Tang, “Action recognition with trajectory-pooled deep-convolutional descriptors,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 4305–4314
2015
Cited alongside, same era.
2015
Cited alongside, same era.
V. Veeriah, N. Zhuang, and G.-J. Qi, “Differential recurrent neural networks for action recognition,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 4041–4049
2015
Cited alongside, same era.
2017
Later among the works it cites.
J. Liu, A. Shahroudy, D. Xu, A. K. Chichung, and G. Wang, “Skeleton-based action recognition using spatio-temporal lstm network with trust gates,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2017
2017
Later among the works it cites.
M. Asadi-Aghbolaghi, H. Bertiche, V. Roig, S. Kasaei, and S. Escalera, “Action recognition from rgb-d data: Comparison and fusion of spatio-temporal handcrafted features and deep strategies,” in Chalearn Workshop on Action, Gesture, and Emotion Recognition: Large Scale Multimodal Gesture Recognition and Real versus Fake expressed emotions (ICCV17) , 2017
2017
Later among the works it cites.
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” arXiv preprint , pp. 1610–02 357, 2017
2017
Later among the works it cites.
M. Liu, H. Liu, and C. Chen, “Enhanced skeleton visualization for view invariant human action recognition,” Pattern Recognition , vol. 68, pp. 346–362, 2017
2017
Later among the works it cites.
Z. Li, K. Gavrilyuk, E. Gavves, M. Jain, and C. G. Snoek, “Videolstm convolves, attends and flows for action recognition,” Computer Vision and Image Understanding , vol. 166, pp. 41–50, 2018
2018
Closest in time.
K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA , 2018, pp. 18–22
2018
Closest in time.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6450–6459
2018
Closest in time.
Y. Zhou, X. Sun, Z.-J. Zha, and W. Zeng, “Mict: Mixed 3d/2d convolutional tube for human action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 449–458
2018
Closest in time.
2018
Closest in time.
C.-Y. Wu, M. Zaheer, H. Hu, R. Manmatha, A. J. Smola, and P. Krähenbühl, “Compressed video action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6026–6035
2018
Closest in time.
2018
Closest in time.
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 305–321
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
P. Wang, W. Li, Z. Gao, C. Tang, and P. O. Ogunbona, “Depth pooling based large-scale 3-d action recognition with convolutional neural networks,” IEEE Transactions on Multimedia , vol. 20, no. 5, pp. 1051–1061, 2018
2018
Closest in time.
2018
Closest in time.
F. Baradel, C. Wolf, J. Mille, and G. W. Taylor, “Glimpse clouds: Human activity recognition from unstructured feature points,” Computer Vision and Pattern Recognition (CVPR)(To appear) , vol. 3, 2018
2018
Closest in time.