Fetching the paper…
Reading the bibliography…
Spatio-temporal action detection is an important and challenging problem in video understanding.
Recognizing human actions: A local SVM approach
Christian Schüldt, Ivan Laptev, and Barbara Caputo · 2004
Earlier work this paper cites.
Actions as space-time shapes
Moshe Blank, Lena Gorelick, Eli Shechtman, Michal Irani, and Ronen Basri · 2005
Earlier work this paper cites.
Action MACH a spatio-temporal maximum average correlation height filter for action recognition
Mikel D. Rodriguez, Javed Ahmed, and Mubarak Shah · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso A. Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Towards understanding action recognition
Hueihan Jhuang, Juergen Gall, Silvia Zuffi, Cordelia Schmid, and Michael J. Black · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Fei-Fei Li · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
Mark Everingham, S. M. Ali Eslami, Luc Van Gool, Christopher K. I. Williams, John M. Winn, and Andrew Zisserman · 2015
Earlier work this paper cites.
Finding action tubes
Georgia Gkioxari and Jitendra Malik · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir D. Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
Limin Wang, Yu Qiao, and Xiaoou Tang · 2015
Earlier work this paper cites.
Learning to track for spatio-temporal action localization
Philippe Weinzaepfel, Zaïd Harchaoui, and Cordelia Schmid · 2015
Earlier work this paper cites.
Object tracking benchmark
Yi Wu, Jongwoo Lim, and Ming-Hsuan Yang · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
SSD: single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott E. Reed, Cheng-Yang Fu, and Alexander C. Berg · 2016
Earlier work this paper cites.
Multi-region two-stream R-CNN for action detection
Xiaojiang Peng and Cordelia Schmid · 2016
Earlier work this paper cites.
Deep learning for detecting multiple space-time action tubes in videos
Suman Saha, Gurkirt Singh, Michael Sapienza, Philip H. S. Torr, and Fabio Cuzzolin · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Cited alongside, same era.
Actionness estimation using hybrid fully convolutional networks
Limin Wang, Yu Qiao, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
Towards weakly-supervised action localization
Philippe Weinzaepfel, Xavier Martin, and Cordelia Schmid · 2016
Cited alongside, same era.
Quo vadis, action recognition? A new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Every moment counts: Dense detailed labeling of actions in complex videos
Serena Yeung, Olga Russakovsky, Ning Jin, Mykhaylo Andriluka, Greg Mori, and Li Fei-Fei · 2018
Later among the works it cites.
Deep layer aggregation
Fisher Yu, Dequan Wang, Evan Shelhamer, and Trevor Darrell · 2018
Later among the works it cites.
Mmdetection: Open mmlab detection toolbox and benchmark
Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin · 2019
Later among the works it cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Later among the works it cites.
Video action transformer network
Rohit Girdhar, João Carreira, Carl Doersch, and Andrew Zisserman · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Federation Internationale de Gymnastique · 2017
Cited alongside, same era.
The ”something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fründ, Peter Yianilos, Moritz Mueller-Freitag, Florian Hoppe, Christian Thurau, Ingo Bax, and Roland Memisevic · 2017
Cited alongside, same era.
The devil is in the tails: Fine-grained classification in the wild
Grant Van Horn and Pietro Perona · 2017
Cited alongside, same era.
Tube convolutional neural network (T-CNN) for action detection in videos
Rui Hou, Chen Chen, and Mubarak Shah · 2017
Cited alongside, same era.
The THUMOS challenge on action recognition for videos ”in the wild”
Haroon Idrees, Amir Roshan Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah · 2017
Cited alongside, same era.
Action tubelet detector for spatio-temporal action localization
Vicky Kalogeiton, Philippe Weinzaepfel, Vittorio Ferrari, and Cordelia Schmid · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie · 2017
Cited alongside, same era.
Okan Köpüklü, Xiangyu Wei, and Gerhard Rigoll · 2019
Later among the works it cites.
BMN: boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen · 2019
Later among the works it cites.
Multi-moments in time: Learning and interpreting models for multi-action video understanding
Mathew Monfort, Kandan Ramakrishnan, Alex Andonian, Barry A. McNamara, Alex Lascelles, Bowen Pan, Quanfu Fan, Dan Gutfreund, Rogério Schmidt Feris, and Aude Oliva · 2019
Later among the works it cites.
Tacnet: Transition-aware context network for spatio-temporal action detection
Lin Song, Shiwei Zhang, Gang Yu, and Hongbin Sun · 2019
Later among the works it cites.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krähenbühl, and Ross B. Girshick · 2019
Later among the works it cites.
STEP: spatio-temporal progressive learning for video action detection
Xitong Yang, Xiaodong Yang, Ming-Yu Liu, Fanyi Xiao, Larry S. Davis, and Jan Kautz · 2019
Later among the works it cites.
Graph convolutional networks for temporal action localization
Runhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan, Yu Rong, Peilin Zhao, and Junzhou Huang · 2019
Later among the works it cites.
HACS: human action clips and segments dataset for recognition and temporal localization
Hang Zhao, Antonio Torralba, Lorenzo Torresani, and Zhicheng Yan · 2019
Later among the works it cites.
Dance with flow: Two-in-one stream action detection
Jiaojiao Zhao and Cees G. M. Snoek · 2019
Later among the works it cites.
Openmmlab’s next generation video understanding toolbox and benchmark
MMAction2 Contributors · 2020
Later among the works it cites.
Fully convolutional online tracking
Yutao Cui, Cheng Jiang, Limin Wang, and Gangshan Wu · 2020
Later among the works it cites.
The ava-kinetics localized human actions video dataset
Ang Li, Meghana Thotakuri, David A. Ross, João Carreira, Alexander Vostrikov, and Andrew Zisserman · 2020
Later among the works it cites.
Actions as moving points
Yixuan Li, Zixu Wang, Limin Wang, and Gangshan Wu · 2020
Later among the works it cites.
Finegym: A hierarchical video dataset for fine-grained action understanding
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin · 2020
Later among the works it cites.
Asynchronous interaction aggregation for action detection
Jiajun Tang, Jin Xia, Xinzhi Mu, Bo Pang, and Cewu Lu · 2020
Later among the works it cites.
Context-aware RCNN: A baseline for action detection in videos
Jianchao Wu, Zhanghui Kuang, Limin Wang, Wayne Zhang, and Gangshan Wu · 2020
Later among the works it cites.
MEVA: A large-scale multiview, multimodal video dataset for activity detection
Kellie Corona, Katie Osterdahl, Roderic Collins, and Anthony Hoogs · 2021
Closest in time.