Fetching the paper…
Reading the bibliography…
Understanding a person's behavior from their 3D motion is a fundamental problem in computer vision with many applications.
Siva, P., Xiang, T.: Weakly supervised action detection. In: BMVC. vol. 2, p. 6. Citeseer (2011)
2011
Earlier work this paper cites.
Bloom, V., Makris, D., Argyriou, V.: G3D: A gaming action dataset and real time action recognition evaluation framework. In: 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops. pp. 7–12. IEEE (2012)
2012
Earlier work this paper cites.
Sung, J., Ponce, C., Selman, B., Saxena, A.: Unstructured human activity detection from RGBD images. 2012 IEEE International Conference on Robotics and Automation pp. 842–849 (2012)
2012
Earlier work this paper cites.
Yun, K., Honorio, J., Chattopadhyay, D., Berg, T.L., Samaras, D.: Two-person interaction detection using body-pose features and multiple instance learning. 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops pp. 28–35 (2012)
2012
Earlier work this paper cites.
Oneata, D., Verbeek, J., Schmid, C.: Action and event recognition with fisher vectors on a compact feature set. In: Proceedings of the IEEE international conference on computer vision. pp. 1817–1824 (2013)
2013
Earlier work this paper cites.
Chen, W., Xiong, C., Xu, R., Corso, J.J.: Actionness ranking with lattice conditional ordinal random fields. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 748–755 (2014)
2014
Earlier work this paper cites.
Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 580–587 (2014)
2014
Earlier work this paper cites.
Jain, M., Van Gemert, J., Jégou, H., Bouthemy, P., Snoek, C.G.: Action localization with tubelets from motion. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 740–747 (2014)
2014
Earlier work this paper cites.
Jiang, Y.G., Liu, J., Roshan Zamir, A., Toderici, G., Laptev, I., Shah, M., Sukthankar, R.: THUMOS challenge: Action recognition with a large number of classes. http://crcv.ucf.edu/THUMOS14/ (2014)
2014
Earlier work this paper cites.
Karaman, S., Seidenari, L., Del Bimbo, A.: Fast saliency based pooling of fisher encoded dense trajectories. In: ECCV THUMOS Workshop. vol. 1, p. 5 (2014)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Lillo, I., Soto, A., Niebles, J.C.: Discriminative hierarchical modeling of spatio-temporally composable human activities. In: 2014 IEEE Conference on Computer Vision and Pattern Recognition. pp. 812–819 (2014). https://doi.org/10.1109/CVPR.2014.109
2014
Earlier work this paper cites.
Wang, L., Qiao, Y., Tang, X.: Action recognition and detection by combining motion and appearance features. THUMOS14 Action Recognition Challenge 1
2014
Earlier work this paper cites.
Gkioxari, G., Malik, J.: Finding action tubes. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 759–768 (2015)
2015
Earlier work this paper cites.
Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., Black, M.J.: SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia) 34
2015
Earlier work this paper cites.
Ramanathan, V., Li, C., Deng, J., Han, W., Li, Z., Gu, K., Song, Y., Bengio, S., Rosenberg, C., Fei-Fei, L.: Learning semantic relationships for better action retrieval in images. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1100–1109 (2015)
2015
Earlier work this paper cites.
Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28
2015
Earlier work this paper cites.
Vedantam, R., Lawrence Zitnick, C., Parikh, D.: CIDEr: Consensus-based image description evaluation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4566–4575 (2015)
2015
Earlier work this paper cites.
Wu, C., Zhang, J., Savarese, S., Saxena, A.: Watch-n-patch: Unsupervised understanding of actions and relations. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 4362–4370 (2015). https://doi.org/10.1109/CVPR.2015.7299065
2015
Earlier work this paper cites.
Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization. arXiv preprint arXiv:1607.06450 (2016)
2016
Earlier work this paper cites.
Escorcia, V., Heilbron, F.C., Niebles, J.C., Ghanem, B.: DAPS: Deep action proposals for action understanding. In: European Conference on Computer Vision. pp. 768–784. Springer (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) pp. 770–778 (2016)
2016
Earlier work this paper cites.
Heilbron, F.C., Niebles, J.C., Ghanem, B.: Fast temporal activity proposals for efficient detection of human actions in untrimmed videos. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1914–1923 (2016)
2016
Earlier work this paper cites.
Li, Y., Lan, C., Xing, J., Zeng, W., Yuan, C., Liu, J.: Online human action detection using joint classification-regression recurrent neural networks. In: European conference on computer vision. pp. 203–220. Springer (2016)
2016
Earlier work this paper cites.
Peng, X., Schmid, C.: Multi-region two-stream R-CNN for action detection. In: European conference on computer vision. pp. 744–759. Springer (2016)
2016
Earlier work this paper cites.
Shou, Z., Wang, D., Chang, S.: Action temporal localization in untrimmed videos via multi-stage CNNs. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition. pp. 1049–1058
2016
Cited alongside, same era.
Stewart, R., Andriluka, M., Ng, A.: End-to-end people detection in crowded scenes. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) pp. 2325–2333 (2016)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? A new model and the Kinetics dataset. In: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6299–6308 (2017)
2017
Cited alongside, same era.
Shi, L., Zhang, Y., Cheng, J., Lu, H.: Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 12026–12035 (2019)
2019
Later among the works it cites.
Sun, C., Myers, A., Vondrick, C., Murphy, K., Schmid, C.: VideoBert: A joint model for video and language representation learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 7464–7473 (2019)
2019
Later among the works it cites.
Zeng, R., Huang, W., Tan, M., Rong, Y., Zhao, P., Huang, J., Gan, C.: Graph convolutional networks for temporal action localization. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 7094–7103 (2019)
2019
Later among the works it cites.
Bai, Y., Wang, Y., Tong, Y., Yang, Y., Liu, Q., Liu, J.: Boundary content graph neural network for temporal action proposal generation. In: European Conference on Computer Vision. pp. 121–137. Springer (2020)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Heilbron, F.C., Barrios, W., Escorcia, V., Ghanem, B.: SCC: Semantic context cascade for efficient action detection. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 3175–3184. IEEE (2017)
2017
Cited alongside, same era.
Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollár, P.: Focal loss for dense object detection. In: Proceedings of the IEEE international conference on computer vision. pp. 2980–2988 (2017)
2017
Cited alongside, same era.
Liu, C., Hu, Y., Li, Y., Song, S., Liu, J.: PKU-MMD: A large scale benchmark for skeleton-based human action understanding. In: Proceedings of the Workshop on Visual Analysis in Smart and Connected Communities. p. 1–8. VSCC ’17, Association for Computing Machinery, New York, NY, USA (2017). https://doi.org/10.1145/3132734.3132739, https://doi.org/10.1145/3132734.3132739
2017
Cited alongside, same era.
Papoutsakis, K., Panagiotakis, C., Argyros, A.A.: Temporal action co-segmentation in 3D motion capture data and videos. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6827–6836 (2017)
2017
Cited alongside, same era.
Shou, Z., Chan, J., Zareian, A., Miyazawa, K., Chang, S.F.: CDC: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 5734–5743 (2017)
2017
Cited alongside, same era.
Singh, G., Saha, S., Sapienza, M., Torr, P.H., Cuzzolin, F.: Online real-time multiple spatiotemporal action localisation and prediction. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 3637–3646 (2017)
2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Advances in neural information processing systems. pp. 5998–6008 (2017)
2017
Cited alongside, same era.
Zhao, Y., Xiong, Y., Wang, L., Wu, Z., Tang, X., Lin, D.: Temporal action detection with structured segment networks. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 2914–2923 (2017)
2017
Cited alongside, same era.
2020
Later among the works it cites.
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T.J., Child, R., Ramesh, A., Ziegler, D.M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., Amodei, D.: Language models are few-shot learners. Advances in Neural Information Processing Systems (2020)
2020
Later among the works it cites.
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: European Conference on Computer Vision. pp. 213–229. Springer (2020)
2020
Later among the works it cites.
Cui, R., Zhu, A., Wu, J., Hua, G.: Skeleton-based attention-aware spatial–temporal model for action detection and recognition. IET Computer Vision 14
2020
Later among the works it cites.
2020
Later among the works it cites.
Kim, D.J., Sun, X., Choi, J., Lin, S., Kweon, I.S.: Detecting human-object interactions with action co-occurrence priors. In: European Conference on Computer Vision. pp. 718–736. Springer (2020)
2020
Later among the works it cites.
Taheri, O., Ghorbani, N., Black, M.J., Tzionas, D.: GRAB: A dataset of whole-body human grasping of objects. In: Computer Vision – ECCV 2020. vol. LNCS 12355, pp. 581–600. Springer International Publishing, Cham (Aug 2020)
2020
Later among the works it cites.
Xia, H., Zhan, Y.: A survey on temporal action localization. IEEE Access 8
2020
Later among the works it cites.
Xu, M., Zhao, C., Rojas, D.S., Thabet, A., Ghanem, B.: G-TAD: Sub-graph localization for temporal action detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10156–10165 (2020)
2020
Later among the works it cites.
Alwassel, H., Giancola, S., Ghanem, B.: TSP: Temporally-sensitive pretraining of video encoders for localization tasks. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 3173–3183 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Lin, C., Xu, C., Luo, D., Wang, Y., Tai, Y., Wang, C., Li, J., Huang, F., Fu, Y.: Learning salient boundary feature for anchor-free temporal action localization. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 3320–3329 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Punnakkal, A.R., Chandrasekaran, A., Athanasiou, N., Quiros-Ramirez, A., Black, M.J.: BABEL: Bodies, action and behavior with english labels. In: Proceedings IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR). pp. 722–731 (Jun 2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Sun, J., Yu, L., Dong, P., Lu, B., Zhou, B.: Adversarial inverse reinforcement learning with self-attention dynamics model. IEEE Robotics and Automation Letters 6
2021
Later among the works it cites.
Tan, J., Tang, J., Wang, L., Wu, G.: Relaxed transformer decoders for direct action proposal generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 13526–13535 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., Dai, J.: Deformable DETR: Deformable transformers for end-to-end object detection. In: International Conference on Learning Representations (2021), https://openreview.net/forum?id=gZ9hCDWe6ke
2021
Later among the works it cites.
Sun, J., Huang, D.A., Lu, B., Liu, Y., Zhou, B., Garg, A.: Plate: Visually-grounded planning with transformers in procedural tasks. IEEE Robotics and Automation Letters pp. 1–1 (2022). https://doi.org/10.1109/LRA.2022.3150855
2022
Closest in time.
Yang, L., Xu, Y., Yuan, C., Liu, W., Li, B., Hu, W.: Improving visual grounding with visual-linguistic verification and iterative reasoning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2022)
2022
Closest in time.