Fetching the paper…
Reading the bibliography…
While deep learning has been widely used for video analytics, such as video classification and action detection, dense action detection with fast-moving subjects from sports videos is still challenging.
Self-supervised learning for semi-supervised temporal action proposal. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1905–1914
Xiang Wang, Shiwei Zhang, Zhiwu Qing, Yuanjie Shao, Changxin Gao, and Nong Sang. 2021 · 1914
Earlier work this paper cites.
A hierarchical deep temporal model for group activity recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1971–1980
Mostafa S Ibrahim, Srikanth Muralidharan, Zhiwei Deng, Arash Vahdat, and Greg Mori. 2016 · 1980
Earlier work this paper cites.
Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05) , Vol. 1. Ieee, 886–893
Navneet Dalal and Bill Triggs. 2005 · 2005
Earlier work this paper cites.
Modeling temporal structure of decomposable motion segments for activity classification. In European conference on computer vision . Springer, 392–405
Juan Carlos Niebles, Chih-Wei Chen, and Li Fei-Fei. 2010 · 2010
Earlier work this paper cites.
Referee Biographical Information
Olympedia. 2012 · 2012
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition . 1725–1732
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. 2014 · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Action recognition in realistic sports videos
Khurram Soomro and Amir R Zamir. 2014 · 2014
Earlier work this paper cites.
Action recognition and detection by combining motion and appearance features
Limin Wang, Yu Qiao, Xiaoou Tang, et al · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 961–970
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision . 4489–4497
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4694–4702
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. 2015 · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. 2016 · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1049–1058
Zheng Shou, Dongang Wang, and Shih-Fu Chang. 2016 · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision . Springer, 20–36
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016 · 2016
Earlier work this paper cites.
End-to-end learning of action detection from frame glimpses in videos. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2678–2687
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Temporal action localization with pyramid of score distribution features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 3093–3102
Jun Yuan, Bingbing Ni, Xiaokang Yang, and Ashraf A Kassim. 2016 · 2016
Earlier work this paper cites.
Am I a baller? basketball performance assessment from first-person videos. In Proceedings of the IEEE international conference on computer vision . 2177–2185
Gedas Bertasius, Hyun Soo Park, Stella X Yu, and Jianbo Shi. 2017 · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6299–6308
Joao Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Temporal context network for activity localization in videos. In Proceedings of the IEEE International Conference on Computer Vision . 5793–5802
Xiyang Dai, Bharat Singh, Guyue Zhang, Larry S Davis, and Yan Qiu Chen. 2017 · 2017
Earlier work this paper cites.
TenniSet: A Dataset for Dense Fine-Grained Event Recognition, Localisation and Description. In 2017 International Conference on Digital Image Computing: Techniques and Applications (DICTA) . IEEE, 1–8
Hayden Faulkner and Anthony Dick. 2017 · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision . 5842–5850
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Deep learning based basketball video analysis for intelligent arena application
Wu Liu, Chenggang Clarence Yan, Jiangyu Liu, and Huadong Ma. 2017 · 2017
Earlier work this paper cites.
Autoaugment: Learning augmentation policies from data
Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le. 2018 · 2018
Cited alongside, same era.
Ava: A video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6047–6056
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition . 6546–6555
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018 · 2018
Cited alongside, same era.
Bsn: Boundary sensitive network for temporal action proposal generation. In Proceedings of the European conference on computer vision (ECCV) . 3–19
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang. 2018 · 2018
Cited alongside, same era.
A comprehensive study of deep video action recognition
Yi Zhu, Xinyu Li, Chunhui Liu, Mohammadreza Zolfaghari, Yuanjun Xiong, Chongruo Wu, Zhi Zhang, Joseph Tighe, R Manmatha, and Mu Li. 2020 · 2020
Later among the works it cites.
Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 6836–6846
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021 · 2021
Later among the works it cites.
Is Space-Time Attention All You Need for Video Understanding?. In International Conference on Machine Learning . PMLR, 813–824
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021 · 2021
Later among the works it cites.
Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4508–4519
Adrien Deliege, Anthony Cioppa, Silvio Giancola, Meisam J Seikavandi, Jacob V Dueholm, Kamal Nasrollahi, Bernard Ghanem, Thomas B Moeslund, and Marc Van Droogenbroeck. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention clusters: Purely attention based local feature integration for video classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 7834–7843
Xiang Long, Chuang Gan, Gerard De Melo, Jiajun Wu, Xiao Liu, and Shilei Wen. 2018 · 2018
Cited alongside, same era.
Sport action recognition with siamese spatio-temporal cnns: Application to table tennis. In 2018 International Conference on Content-Based Multimedia Indexing (CBMI) . IEEE, 1–6
Pierre-Etienne Martin, Jenny Benois-Pineau, Renaud Péteri, and Julien Morlier. 2018 · 2018
Cited alongside, same era.
Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 7794–7803
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018 · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In Proceedings of the European conference on computer vision (ECCV) . 305–321
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. 2018 · 2018
Cited alongside, same era.
Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision . 6202–6211
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019 · 2019
Cited alongside, same era.
Use what you have: Video retrieval using representations from collaborative experts
Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. 2019a · 2019
Cited alongside, same era.
Building effective short video recommendation. In 2019 IEEE International Conference on Multimedia & Expo Workshops (ICMEW) . IEEE, 651–656
Yang Liu, Cheng Lyu, Zhiyuan Liu, and Dacheng Tao. 2019b · 2019
Cited alongside, same era.
PaddlePaddle: An open-source deep learning platform from industrial practice
Yanjun Ma, Dianhai Yu, Tian Wu, and Haifeng Wang. 2019 · 2019
Cited alongside, same era.
Dual Encoding for Video Retrieval by Text
Jianfeng Dong, Xirong Li, Chaoxi Xu, Xun Yang, Gang Yang, Xun Wang, and Meng Wang. 2022 · 2021
Later among the works it cites.
Understanding test-time augmentation. In International Conference on Neural Information Processing . Springer, 558–569
Masanari Kimura. 2021 · 2021
Later among the works it cites.
Movinets: Mobile video networks for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16020–16030
Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong. 2021 · 2021
Later among the works it cites.
Contrastive learning for sports video: Unsupervised player classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4528–4536
Maria Koshkina, Hemanth Pidaparthy, and James H Elder. 2021 · 2021
Later among the works it cites.
Table tennis stroke recognition using two-dimensional human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4576–4584
Kaustubh Milind Kulkarni and Sucheth Shenoy. 2021 · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 10012–10022
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021 · 2021
Later among the works it cites.
Three-Stream 3D/1D CNN for Fine-Grained Action Classification and Segmentation in Table Tennis. In Proceedings of the 4th International Workshop on Multimedia Content Analysis in Sports . 35–41
Pierre-Etienne Martin, Jenny Benois-Pineau, Renaud Péteri, and Julien Morlier. 2021 · 2021
Later among the works it cites.
Activity graph transformer for temporal action localization
Megha Nawhal and Greg Mori. 2021 · 2021
Later among the works it cites.
Temporal context aggregation network for temporal action proposal refinement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 485–494
Zhiwu Qing, Haisheng Su, Weihao Gan, Dongliang Wang, Wei Wu, Xiang Wang, Yu Qiao, Junjie Yan, Changxin Gao, and Nong Sang. 2021 · 2021
Later among the works it cites.
Toward the Perfect Stroke: A Multimodal Approach for Table Tennis Stroke Evaluation. In 2021 Thirteenth International Conference on Mobile Computing and Ubiquitous Network (ICMU) . IEEE, 1–5
Panyawut Sri-Iesaranusorn, Felan Carlo Garcia, Francis Tiausas, Supatsara Wattanakriengkrai, Kazushi Ikeda, and Junichiro Yoshimoto. 2021 · 2021
Later among the works it cites.
Bsn++: Complementary boundary regressor with scale-balanced relation modeling for temporal action proposal generation. In Proceedings of the AAAI conference on artificial intelligence , Vol. 35. 2602–2610
Haisheng Su, Weihao Gan, Wei Wu, Yu Qiao, and Junjie Yan. 2021 · 2021
Later among the works it cites.
Boundary-sensitive pre-training for temporal localization in videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7220–7230
Mengmeng Xu, Juan-Manuel Pérez-Rúa, Victor Escorcia, Brais Martinez, Xiatian Zhu, Li Zhang, Bernard Ghanem, and Tao Xiang. 2021 · 2021
Later among the works it cites.
Machine learning in real-time internet of things (iot) systems: A survey
Jiang Bian, Abdullah Al Arafat, Haoyi Xiong, Jing Li, Li Li, Hongyang Chen, Jun Wang, Dejing Dou, and Zhishan Guo. 2022 · 2022
Closest in time.
Heterogeneous Graph Contrastive Learning Network for Personalized Micro-Video Recommendation
Desheng Cai, Shengsheng Qian, Quan Fang, Jun Hu, Wenkui Ding, and Changsheng Xu. 2023 · 2022
Closest in time.
Video swin transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 3202–3211
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. 2022 · 2022
Closest in time.
Pose is all you need: The pose only group activity recognition system (pogars)
Haritha Thilakarathne, Aiden Nibali, Zhen He, and Stuart Morgan. 2022 · 2022
Closest in time.
Masked feature prediction for self-supervised visual pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14668–14678
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. 2022 · 2022
Closest in time.
A Survey on Video Action Recognition in Sports: Datasets, Methods and Applications
Fei Wu, Qingzhong Wang, Jiang Bian, Ning Ding, Feixiang Lu, Jun Cheng, Dejing Dou, and Haoyi Xiong. 2023 · 2022
Closest in time.
Classification of seed corn ears based on custom lightweight convolutional neural network and improved training strategies
Xiang Ma, Yonglei Li, Lipengcheng Wan, Zexin Xu, Jiannong Song, and Jinqiu Huang. 2023 · 2023
Closest in time.
Videomae v2: Scaling video masked autoencoders with dual masking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14549–14560
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. 2023 · 2023
Closest in time.