Fetching the paper…
Reading the bibliography…
We introduce Activity Graph Transformer, an end-to-end learnable model for temporal action localization, that receives a video as input and directly predicts a set of action instances that appear in the video.
On space-time interest points
Ivan Laptev · 2005
Earlier work this paper cites.
Human detection using oriented histograms of flow and appearance
Navneet Dalal, Bill Triggs, and Cordelia Schmid · 2006
Earlier work this paper cites.
Temporal localization of actions with actoms
Adrien Gaidon, Zaid Harchaoui, and Cordelia Schmid · 2013
Earlier work this paper cites.
Action and event recognition with fisher vectors on a compact feature set
Dan Oneata, Jakob Verbeek, and Cordelia Schmid · 2013
Earlier work this paper cites.
Combining the right features for complex event recognition
Kevin Tang, Bangpeng Yao, Li Fei-Fei, and Daphne Koller · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Action localization with tubelets from motion
Mihir Jain, Jan Van Gemert, Hervé Jégou, Patrick Bouthemy, and Cees GM Snoek · 2014
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Action recognition and detection by combining motion and appearance features
Limin Wang, Yu Qiao, and Xiaoou Tang · 2014
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Finding action tubes
Georgia Gkioxari and Jitendra Malik · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
Limin Wang, Yu Qiao, and Xiaoou Tang · 2015
Earlier work this paper cites.
Fast temporal activity proposals for efficient detection of human actions in untrimmed videos
Fabian Caba Heilbron, Juan Carlos Niebles, and Bernard Ghanem · 2016
Earlier work this paper cites.
Daps: Deep action proposals for action understanding
Victor Escorcia, Fabian Caba Heilbron, Juan Carlos Niebles, and Bernard Ghanem · 2016
Earlier work this paper cites.
Structural-rnn: Deep learning on spatio-temporal graphs
Ashesh Jain, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena · 2016
Earlier work this paper cites.
Learning activity progression in lstms for activity detection and early detection
Shugao Ma, Leonid Sigal, and Stan Sclaroff · 2016
Earlier work this paper cites.
Temporal action detection using a statistical language model
Alexander Richard and Juergen Gall · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
A multi-stream bi-directional recurrent neural network for fine-grained action detection
Bharat Singh, Tim K Marks, Michael Jones, Oncel Tuzel, and Ming Shao · 2016
Earlier work this paper cites.
End-to-end people detection in crowded scenes
Russell Stewart, Mykhaylo Andriluka, and Andrew Y Ng · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2016
Cited alongside, same era.
Temporal action localization with pyramid of score distribution features
Jun Yuan, Bingbing Ni, Xiaokang Yang, and Ashraf A Kassim · 2016
Cited alongside, same era.
Real-time action recognition with enhanced motion vector cnns
Bowen Zhang, Limin Wang, Zhe Wang, Yu Qiao, and Hanli Wang · 2016
Cited alongside, same era.
End-to-end, single-stream temporal action detection in untrimmed videos
Rethinking the faster r-cnn architecture for temporal action localization
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A Ross, Jia Deng, and Rahul Sukthankar · 2018
Later among the works it cites.
Bsn: Boundary sensitive network for temporal action proposal generation
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang · 2018
Later among the works it cites.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Łukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Later among the works it cites.
Learning latent super-events to detect multiple activities in videos
AJ Piergiovanni and Michael S Ryoo · 2018
Later among the works it cites.
Autoloc: Weakly-supervised temporal action localization in untrimmed videos
Zheng Shou, Hang Gao, Lei Zhang, Kazuyuki Miyazawa, and Shih-Fu Chang · 2018
Later among the works it cites.
Graph attention networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shyamal Buch, Victor Escorcia, Bernard Ghanem, Li Fei-Fei, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Sst: Single-stream temporal action proposals
Shyamal Buch, Victor Escorcia, Chuanqi Shen, Bernard Ghanem, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Temporal context network for activity localization in videos
Xiyang Dai, Bharat Singh, Guyue Zhang, Larry S Davis, and Yan Qiu Chen · 2017
Cited alongside, same era.
Predictive-corrective networks for action detection
Achal Dave, Olga Russakovsky, and Deva Ramanan · 2017
Cited alongside, same era.
Turn tap: Temporal unit regression network for temporal action proposals
Jiyang Gao, Zhenheng Yang, Kan Chen, Chen Sun, and Ram Nevatia · 2017
Cited alongside, same era.
Scc: Semantic context cascade for efficient action detection
Fabian Caba Heilbron, Wayner Barrios, Victor Escorcia, and Bernard Ghanem · 2017
Cited alongside, same era.
Later among the works it cites.
Videos as space-time region graphs
Xiaolong Wang and Abhinav Gupta · 2018
Later among the works it cites.
Video action transformer network
Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019
Later among the works it cites.
Videograph: Recognizing minutes-long human activities in videos
Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders · 2019
Later among the works it cites.
Bmn: Boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen · 2019
Later among the works it cites.
Temporal gaussian mixture layer for videos
AJ Piergiovanni and Michael Ryoo · 2019
Later among the works it cites.
Video relationship reasoning using gated spatio-temporal energy graph
Yao-Hung Hubert Tsai, Santosh Divvala, Louis-Philippe Morency, Ruslan Salakhutdinov, and Ali Farhadi · 2019
Later among the works it cites.
Graph convolutional networks for temporal action localization
Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan · 2019
Later among the works it cites.
Boundary content graph neural network for temporal action proposal generation
Yueran Bai, Yingying Wang, Yunhai Tong, Yang Yang, Qiyue Liu, and Junhui Liu · 2020
Later among the works it cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, , Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2020
Later among the works it cites.
Actionbytes: Learning from trimmed videos to localize actions
Mihir Jain, Amir Ghodrati, and Cees GM Snoek · 2020
Later among the works it cites.
Representation learning on visual-symbolic graphs for video understanding
Effrosyni Mavroudi, Benjamín Béjar Haro, and René Vidal · 2020
Later among the works it cites.
Ego-topo: Environment affordances from egocentric video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman · 2020
Later among the works it cites.
Spatio-temporal graph for video captioning with knowledge distillation
Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, and Juan Carlos Niebles · 2020
Later among the works it cites.
Avid dataset: Anonymized videos from diverse countries
AJ Piergiovanni and Michael S Ryoo · 2020
Later among the works it cites.
G-tad: Sub-graph localization for temporal action detection
Mengmeng Xu, Chen Zhao, David S Rojas, Ali Thabet, and Bernard Ghanem · 2020
Later among the works it cites.
Bottom-up temporal action localization with mutual regularization
Peisen Zhao, Lingxi Xie, Chen Ju, Ya Zhang, Yanfeng Wang, and Qi Tian · 2020
Later among the works it cites.