Fetching the paper…
Reading the bibliography…
Traditional temporal action detection (TAD) usually handles untrimmed videos with small number of action instances from a single label (e.g., ActivityNet, THUMOS).
Keypoint-based keyframe selection
Genliang Guan, Zhiyong Wang, Shiyang Lu, Jeremiah Da Deng, and David Dagan Feng · 2013
Earlier work this paper cites.
Action and event recognition with fisher vectors on a compact feature set
Dan Oneata, Jakob Verbeek, and Cordelia Schmid · 2013
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Bag-of-fragments: Selecting and encoding video fragments for event detection and recounting
Pascal Mettes, Jan C. van Gemert, Spencer Cappallo, Thomas Mensink, and Cees G. M. Snoek · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Deformable convolutional networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei · 2017
Earlier work this paper cites.
The THUMOS challenge on action recognition for videos "in the wild"
Haroon Idrees, Amir Roshan Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, João Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie · 2017
Earlier work this paper cites.
UntrimmedNets for weakly supervised action recognition and detection
Limin Wang, Yuanjun Xiong, Dahua Lin, and Luc Van Gool · 2017
Earlier work this paper cites.
R-C3D: region convolutional 3d network for temporal activity detection
Huijuan Xu, Abir Das, and Kate Saenko · 2017
Earlier work this paper cites.
Cascade R-CNN: delving into high quality object detection
Zhaowei Cai and Nuno Vasconcelos · 2018
Earlier work this paper cites.
Rethinking the faster R-CNN architecture for temporal action localization
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A. Ross, Jia Deng, and Rahul Sukthankar · 2018
Earlier work this paper cites.
BSN: boundary sensitive network for temporal action proposal generation
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang · 2018
Earlier work this paper cites.
Learning latent super-events to detect multiple activities in videos
A. J. Piergiovanni and Michael S. Ryoo · 2018
Earlier work this paper cites.
Every moment counts: Dense detailed labeling of actions in complex videos
Serena Yeung, Olga Russakovsky, Ning Jin, Mykhaylo Andriluka, Greg Mori, and Li Fei-Fei · 2018
Earlier work this paper cites.
BMN: boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen · 2019
Cited alongside, same era.
Temporal gaussian mixture layer for videos
A. J. Piergiovanni and Michael S. Ryoo · 2019
Cited alongside, same era.
Fast and robust dynamic hand gesture recognition via key frames extraction and feature fusion
Hao Tang, Hong Liu, Wei Xiao, and Nicu Sebe · 2019
Cited alongside, same era.
FCOS: fully convolutional one-stage object detection
Zhi Tian, Chunhua Shen, Hao Chen, and Tong He · 2019
Cited alongside, same era.
Reppoints: Point set representation for object detection
Ze Yang, Shaohui Liu, Han Hu, Liwei Wang, and Stephen Lin · 2019
Cited alongside, same era.
Graph convolutional networks for temporal action localization
Runhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan, Yu Rong, Peilin Zhao, and Junzhou Huang · 2019
Learning salient boundary feature for anchor-free temporal action localization
Chuming Lin, Chengming Xu, Donghao Luo, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yanwei Fu · 2021
Later among the works it cites.
Activity graph transformer for temporal action localization
Megha Nawhal and Greg Mori · 2021
Later among the works it cites.
Temporal context aggregation network for temporal action proposal refinement
Zhiwu Qing, Haisheng Su, Weihao Gan, Dongliang Wang, Wei Wu, Xiang Wang, Yu Qiao, Junjie Yan, Changxin Gao, and Nong Sang · 2021
Later among the works it cites.
BSN++: complementary boundary regressor with scale-balanced relation modeling for temporal action proposal generation
Haisheng Su, Weihao Gan, Wei Wu, Yu Qiao, and Junjie Yan · 2021
Later among the works it cites.
Sparse R-CNN: end-to-end object detection with learnable proposals
Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, and Ping Luo · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
HACS: human action clips and segments dataset for recognition and temporal localization
Hang Zhao, Antonio Torralba, Lorenzo Torresani, and Zhicheng Yan · 2019
Cited alongside, same era.
Deformable convnets V2: more deformable, better results
Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Accurate temporal action proposal generation with relation-aware pyramid network
Jialin Gao, Zhixiang Shi, Guanshuo Wang, Jiani Li, Yufeng Yuan, Shiming Ge, and Xi Zhou · 2020
Cited alongside, same era.
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick · 2020
Cited alongside, same era.
Actions as moving points
Yixuan Li, Zixu Wang, Limin Wang, and Gangshan Wu · 2020
Cited alongside, same era.
Later among the works it cites.
Relaxed transformer decoders for direct action proposal generation
Jing Tan, Jiaqi Tang, Limin Wang, and Gangshan Wu · 2021
Later among the works it cites.
Modeling multi-label action dependencies for temporal action localization
Praveen Tirupattur, Kevin Duarte, Yogesh S. Rawat, and Mubarak Shah · 2021
Later among the works it cites.
TDN: temporal difference networks for efficient action recognition
Limin Wang, Zhan Tong, Bin Ji, and Gangshan Wu · 2021
Later among the works it cites.
Towards high-quality temporal action detection with sparse proposals
Jiannan Wu, Peize Sun, Shoufa Chen, Jiewen Yang, Zihao Qi, Lan Ma, and Ping Luo · 2021
Later among the works it cites.
Deformable DETR: deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2021
Later among the works it cites.
Enriching local and global contexts for temporal action localization
Zixin Zhu, Wei Tang, Le Wang, Nanning Zheng, and Gang Hua · 2021
Later among the works it cites.
DCAN: improving temporal action detection via dual context aggregation
Guo Chen, Yin-Dong Zheng, Limin Wang, and Tong Lu · 2022
Closest in time.
MS-TCT: multi-scale temporal convtransformer for action detection
Rui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo, and François Brémond · 2022
Closest in time.
AdaMixer: A fast-converging query-based object detector
Ziteng Gao, Limin Wang, Bing Han, and Sheng Guo · 2022
Closest in time.
An empirical study of end-to-end temporal action detection
Xiaolong Liu, Song Bai, and Xiang Bai · 2022
Closest in time.
End-to-end temporal action detection with transformer
Xiaolong Liu, Qimeng Wang, Yao Hu, Xu Tang, Song Bai, and Xiang Bai · 2022
Closest in time.
Learning strides in convolutional neural networks
Rachid Riad, Olivier Teboul, David Grangier, and Neil Zeghidour · 2022
Closest in time.
BasicTAD: an astounding rgb-only baseline for temporal action detection
Min Yang, Guo Chen, Yin-Dong Zheng, Tong Lu, and Limin Wang · 2022
Closest in time.