Fetching the paper…
Reading the bibliography…
In this paper, we propose a \textbf{Tr}ansformer-based RGB-D \textbf{e}gocentric \textbf{a}ction \textbf{r}ecognition framework, called Trear.
M. Ma, H. Fan, and K. M. Kitani, “Going deeper into first-person activity recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 1894–1903
1903
Earlier work this paper cites.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 1933–1941
1941
Earlier work this paper cites.
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld, “Learning realistic human actions from movies,” in 2008 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2008, pp. 1–8
2008
Earlier work this paper cites.
H. Wang, M. M. Ullah, A. Klaser, I. Laptev, and C. Schmid, “Evaluation of local spatio-temporal features for action recognition,” in BMCV , 2009
2009
Earlier work this paper cites.
A. Fathi, A. Farhadi, and J. M. Rehg, “Understanding egocentric activities,” in 2011 international conference on computer vision . IEEE, 2011, pp. 407–414
2011
Earlier work this paper cites.
O. Oreifej and Z. Liu, “Hon4d: Histogram of oriented 4d normals for activity recognition from depth sequences,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2013, pp. 716–723
2013
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in Advances in neural information processing systems , 2014, pp. 568–576
2014
Earlier work this paper cites.
M. Moghimi, P. Azagra, L. Montesano, A. C. Murillo, and S. Belongie, “Experiments on an rgb-d wearable vision system for egocentric activity recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , 2014, pp. 597–603
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
E. Ohn-Bar and M. M. Trivedi, “Hand gesture recognition in real time for automotive interfaces: A multimodal vision-based approach and evaluations,” IEEE transactions on intelligent transportation systems , vol. 15, no. 6, pp. 2368–2377, 2014
2014
Earlier work this paper cites.
Y. Kong and Y. Fu, “Bilinear heterogeneous information machine for rgb-d action recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1054–1062
2015
Earlier work this paper cites.
Y. Li, Z. Ye, and J. M. Rehg, “Delving into egocentric actions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 287–295
2015
Earlier work this paper cites.
S. Singh, C. Arora, and C. Jawahar, “First person action recognition using deep learned descriptors,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 2620–2628
2016
Cited alongside, same era.
2016
Cited alongside, same era.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good practices for deep action recognition,” in European conference on computer vision . Springer, 2016, pp. 20–36
2016
Cited alongside, same era.
X. Zhang, Y. Wang, M. Gou, M. Sznaier, and O. Camps, “Efficient temporal sequence comparison and classification using gram matrix embeddings on a riemannian manifold,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 4498–4507
2016
Cited alongside, same era.
H. Lee, M. Jung, and J. Tani, “Recognition of visually perceived compositional human actions by multiple spatio-temporal scales recurrent neural networks,” IEEE Transactions on Cognitive and Developmental Systems , vol. 10, no. 4, pp. 1058–1069, 2018
2018
Later among the works it cites.
S. A. W. Talha, M. Hammouche, E. Ghorbel, A. Fleury, and S. Ambellouis, “Features and classification schemes for view-invariant and real-time human action recognition,” IEEE Transactions on Cognitive and Developmental Systems , vol. 10, no. 4, pp. 894–902, 2018
2018
Later among the works it cites.
P. Wang, W. Li, J. Wan, P. Ogunbona, and X. Liu, “Cooperative training of deep aggregation networks for rgb-d action recognition,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
Y. Tang, Z. Wang, J. Lu, J. Feng, and J. Zhou, “Multi-stream deep neural networks for rgb-d egocentric action recognition,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 29, no. 10, pp. 3001–3015, 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. K. Vishwakarma and K. Singh, “Human activity recognition based on spatial distribution of gradients at sublevels of average energy silhouette images,” IEEE Transactions on Cognitive and Developmental Systems , vol. 9, no. 4, pp. 316–327, 2017
2017
Cited alongside, same era.
P. Wang, W. Li, Z. Gao, Y. Zhang, C. Tang, and P. Ogunbona, “Scene flow to action map: A new representation for rgb-d based action recognition with convolutional neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 595–604
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Cited alongside, same era.
J. Liu and S. Amir and D. Xu and A. Kot and G. Wang,“Skeleton-based action recognition using spatio-temporal lstm network with trust gates”, in IEEE transactions on pattern analysis and machine intelligence , 2017, pp. 3007–3021
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
S. Yan, J. S. Smith, W. Lu, and B. Zhang, “Multibranch attention networks for action recognition in still images,” IEEE Transactions on Cognitive and Developmental Systems , vol. 10, no. 4, pp. 1116–1125, 2018
2018
Cited alongside, same era.
2018
Later among the works it cites.
G. Garcia-Hernando, S. Yuan, S. Baek, and T.-K. Kim, “First-person hand action benchmark with rgb-d videos and 3d hand pose annotations,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 409–419
2018
Later among the works it cites.
P. Wang, W. Li, P. Ogunbona, J. Wan, and S. Escalera, “Rgb-d-based human motion recognition with deep learning: A survey,” Computer Vision and Image Understanding , vol. 171, pp. 118–139, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Sudhakaran, S. Escalera, and O. Lanz, “Lsta: Long short-term attention for egocentric action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 9954–9963
2019
Later among the works it cites.
R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman, “Video action transformer network,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 244–253
2019
Later among the works it cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in Advances in Neural Information Processing Systems , 2019, pp. 13–23
2019
Later among the works it cites.
B. Tekin, F. Bogo, and M. Pollefeys, “H+ o: Unified egocentric recognition of 3d hand-object poses and interactions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 4511–4520
2019
Later among the works it cites.