Fetching the paper…
Reading the bibliography…
Many attempts have been made towards combining RGB and 3D poses for the recognition of Activities of Daily Living (ADL).
1905
Earlier work this paper cites.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” in Computer Vision and Pattern Recognition (CVPR), 2016 IEEE Conference on . IEEE, 2016, pp. 1933–1941
1941
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR09 , 2009
2009
Earlier work this paper cites.
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “Hmdb: a large video database for human motion recognition,” in 2011 International Conference on Computer Vision . IEEE, 2011, pp. 2556–2563
2011
Earlier work this paper cites.
J. Shotton, A. Fitzgibbon, M. Cook, T. Sharp, M. Finocchio, R. Moore, A. Kipman, and A. Blake, “Real-time human pose recognition in parts from single depth images,” in CVPR , 2011
2011
Earlier work this paper cites.
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu, “Action Recognition by Dense Trajectories,” in IEEE Conference on Computer Vision & Pattern Recognition , Colorado Springs, United States, Jun. 2011, pp. 3169–3176. [Online]. Available: http://hal.inria.fr/inria-00583818/en
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in Advances in neural information processing systems , 2014, pp. 568–576
2014
Earlier work this paper cites.
A. Shahroudy, G. Wang, and T.-T. Ng, “Multi-modal feature fusion for action recognition in rgb-d sequences,” in 2014 6th International Symposium on Communications, Control and Signal Processing (ISCCSP) , May 2014, pp. 1–4
2014
Earlier work this paper cites.
J. Wang, X. Nie, Y. Xia, Y. Wu, and S.-C. Zhu, “Cross-view action modeling, learning, and recognition,” in 2014 IEEE Conference on Computer Vision and Pattern Recognition , June 2014, pp. 2649–2656
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res. , vol. 15, no. 1, pp. 1929–1958, Jan. 2014. [Online]. Available: http://dl.acm.org/citation.cfm?id=2627435.2670313
2014
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV) , ser. ICCV ’15. Washington, DC, USA: IEEE Computer Society, 2015, pp. 4489–4497. [Online]. Available: http://dx.doi.org/10.1109/ICCV.2015.510
2015
Earlier work this paper cites.
Y. Du, Y. Fu, and L. Wang, “Skeleton based action recognition with convolutional neural network,” in 2015 3rd IAPR Asian Conference on Pattern Recognition (ACPR) , Nov 2015, pp. 579–583
2015
Earlier work this paper cites.
H. Rahmani and A. Mian, “Learning a non-linear knowledge transfer model for cross-view action recognition,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2015, pp. 2458–2466
2015
Earlier work this paper cites.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Val Gool, “Temporal segment networks: Towards good practices for deep action recognition,” in ECCV , 2016
2016
Earlier work this paper cites.
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+d: A large scale dataset for 3d human activity analysis,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2016
2016
Earlier work this paper cites.
J. Liu, A. Shahroudy, D. Xu, and G. Wang, “Spatio-temporal lstm with trust gates for 3d human action recognition,” in Computer Vision – ECCV 2016 , B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds. Cham: Springer International Publishing, 2016, pp. 816–833
2016
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in Proceedings of the 30th International Conference on Neural Information Processing Systems , ser. NIPS’16. Red Hook, NY, USA: Curran Associates Inc., 2016, p. 892–900
2016
Earlier work this paper cites.
J. Hoffman, S. Gupta, and T. Darrell, “Learning with side information through modality hallucination,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 826–834
2016
Earlier work this paper cites.
J. Hoffman, S. Gupta, J. Leong, S. Guadarrama, and T. Darrell, “Cross-modal adaptation for rgb-d detection,” in 2016 IEEE International Conference on Robotics and Automation (ICRA) , 2016, pp. 5032–5039
2016
Earlier work this paper cites.
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui, “Jointly modeling embedding and translation to bridge video and language,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2016
2016
Earlier work this paper cites.
B. Mahasseni and S. Todorovic, “Regularizing long short term memory with 3d human-skeleton sequences for action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 3054–3062
2016
Earlier work this paper cites.
H. Rahmani and A. Mian., “3d action recognition from novel viewpoints,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2016, pp. 1506–1515
2016
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2017, pp. 4724–4733
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Rahmani and M. Bennamoun, “Learning action recognition model from depth and skeleton videos,” in 2017 IEEE International Conference on Computer Vision (ICCV) , Oct 2017, pp. 5833–5842
2017
Earlier work this paper cites.
F. Baradel, C. Wolf, and J. Mille, “Human action recognition: Pose-based attention draws focus to hands,” in proceedings of the IEEE International Conference on Computer Vision WorkshopS , 2017, pp. 604–613
2017
Cited alongside, same era.
Z. Qiu, T. Yao, and T. Mei, “Learning spatio-temporal representation with pseudo-3d residual networks,” in 2017 IEEE International Conference on Computer Vision (ICCV) . IEEE, 2017, pp. 5534–5542
2017
Cited alongside, same era.
S. Zhang, X. Liu, and J. Xiao, “On geometric features for skeleton-based action recognition using multilayer lstm networks,” in 2017 IEEE Winter Conference on Applications of Computer Vision (WACV) , March 2017, pp. 148–157
2017
Cited alongside, same era.
P. Zhang, C. Lan, J. Xing, W. Zeng, J. Xue, and N. Zheng, “View adaptive recurrent neural networks for high performance human action recognition from skeleton data,” in The IEEE International Conference on Computer Vision (ICCV) , Oct 2017
2017
Cited alongside, same era.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in The IEEE International Conference on Computer Vision (ICCV) , October 2019
2019
Later among the works it cites.
S. Das, R. Dai, M. Koperski, L. Minciullo, L. Garattoni, F. Bremond, and G. Francesca, “Toyota smarthome: Real-world activities of daily living,” in ICCV , 2019
2019
Later among the works it cites.
J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, “Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2019
2019
Later among the works it cites.
L. Shi, Y. Zhang, J. Cheng, and H. Lu, “Skeleton-based action recognition with directed graph neural networks,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Song, C. Lan, J. Xing, W. Zeng, and J. Liu, “An end-to-end spatio-temporal attention model for human action recognition from skeleton data,” in AAAI Conference on Artificial Intelligence , 2017, pp. 4263–4270
2017
Cited alongside, same era.
M. Liu, H. Liu, and C. Chen, “Enhanced skeleton visualization for view invariant human action recognition,” Pattern Recognition , vol. 68, pp. 346 – 362, 2017. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S0031320317300936
2017
Cited alongside, same era.
H.-S. Fang, S. Xie, Y.-W. Tai, and C. Lu, “RMPE: Regional multi-person pose estimation,” in ICCV , 2017
2017
Cited alongside, same era.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in 2017 IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 618–626
2017
Cited alongside, same era.
J. Liu, G. Wang, P. Hu, L.-Y. Duan, and A. C. Kot, “Global context-aware attention lstm networks for 3d action recognition,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , July 2017, pp. 3671–3680
2017
Cited alongside, same era.
I. Lee, D. Kim, S. Kang, and S. Lee, “Ensemble deep learning for skeleton-based action recognition using temporal sliding lstm networks,” in Proceedings of the IEEE International Conference on Computer Vision , 2017
2017
Cited alongside, same era.
X. Wang, R. B. Girshick, A. Gupta, and K. He, “Non-local neural networks,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 7794–7803, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Lei, Z. Yifan, C. Jian, and L. Hanqing, “Two-stream adaptive graph convolutional networks for skeleton-based action recognition,” in CVPR , 2019
2019
Later among the works it cites.
G. Liu, J. Qian, F. Wen, X. Zhu, R. Ying, and P. Liu, “Action recognition based on 3d skeleton and rgb frame fusion,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , Nov 2019, pp. 258–264
2019
Later among the works it cites.
J.-M. Perez-Rua, V. Vielzeuf, S. Pateux, M. Baccouche, and F. Jurie, “Mfas: Multimodal fusion architecture search,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2019
2019
Later among the works it cites.
S. Das, A. Chaudhary, F. Bremond, and M. Thonnat, “Where to focus on for human action recognition?” in 2019 IEEE Winter Conference on Applications of Computer Vision (WACV) , Jan 2019, pp. 71–80
2019
Later among the works it cites.
D. Tran, H. Wang, L. Torresani, and M. Feiszli, “Video classification with channel-separated convolutional networks,” in The IEEE International Conference on Computer Vision (ICCV) , October 2019
2019
Later among the works it cites.
S. Das, M. Thonnat, kaustubh Sakhalkar, M. Koperski, F. Brémond, and G. Francesca, “A new hybrid architecture for human activity recognition from rgb-d videos,” MultiMedia Modeling. MMM 2019 , 2019
2019
Later among the works it cites.
N. Crasto, P. Weinzaepfel, K. Alahari, and C. Schmid, “MARS: Motion-Augmented RGB Stream for Action Recognition,” in CVPR , 2019
2019
Later among the works it cites.
Z. Cao, G. Hidalgo Martinez, T. Simon, S.-E. Wei, and Y. Sheikh, “Openpose: Realtime multi-person 2d pose estimation using part affinity fields,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2019
2019
Later among the works it cites.
L. Shi, Y. Zhang, J. Cheng, and H. Lu, “Skeleton-based action recognition with multi-stream adaptive graph convolutional networks,” IEEE Transactions on Image Processing , vol. 29, pp. 9532–9545, 2020
2020
Later among the works it cites.
S. Das, S. Sharma, R. Dai, F. Bremond, and M. Thonnat, “Vpn: Learning video-pose embedding for activities of daily living,” 2020
2020
Later among the works it cites.
C. Feichtenhofer, “X3d: Expanding architectures for efficient video recognition,” 2020
2020
Later among the works it cites.
M. S. Ryoo, A. Piergiovanni, M. Tan, and A. Angelova, “Assemblenet: Searching for multi-stream neural connectivity in video architectures,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=SJgMK64Ywr
2020
Later among the works it cites.
M. S. Ryoo, A. Piergiovanni, J. Kangaspunta, and A. Angelova, “Assemblenet++: Assembling modality representations via attention connections,” in ECCV , 2020
2020
Later among the works it cites.
A. Piergiovanni, A. Angelova, and M. S. Ryoo, “Evolving losses for unsupervised video representation learning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Later among the works it cites.
Y. Tian, D. Krishnan, and P. Isola, “Contrastive representation distillation,” in International Conference on Learning Representations , 2020
2020
Later among the works it cites.
Q. Guo, X. Wang, Y. Wu, Z. Yu, D. Liang, X. Hu, and P. Luo, “Online knowledge distillation via collaborative learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Later among the works it cites.
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman, “End-to-End Learning of Visual Representations from Uncurated Instructional Videos,” in CVPR , 2020
2020
Later among the works it cites.
D. Yang, R. Dai, Y. Wang, R. Mallick, L. Minciullo, G. Francesca, and F. Bremond, “Selective spatio-temporal aggregation based pose refinement system: Towards understanding human activities in real-world videos,” 2020
2020
Later among the works it cites.
Z. Liu, H. Zhang, Z. Chen, Z. Wang, and W. Ouyang, “Disentangling and unifying graph convolutions for skeleton-based action recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 143–152
2020
Later among the works it cites.
P. Zhang, C. Lan, W. Zeng, J. Xing, J. Xue, and N. Zheng, “Semantics-guided neural networks for efficient skeleton-based human action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2020
2020
Later among the works it cites.
S. Das, M. Thonnat, and F. Bremond, “Looking deeper into time for activities of daily living recognition,” in 2020 IEEE Winter Conference on Applications of Computer Vision (WACV) , 2020, pp. 487–496
2020
Later among the works it cites.