Fetching the paper…
Reading the bibliography…
Weakly-supervised temporal action localization is a problem of learning an action localization model with only video-level action labeling available.
On space-time interest points
Ivan Laptev · 2005
Earlier work this paper cites.
A generic framework of user attention model and its application in video summarization
Yu-Fei Ma, Xian-Sheng Hua, Lie Lu, and Hong-Jiang Zhang · 2005
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Action and event recognition with fisher vectors on a compact feature set
Dan Oneata, Jakob Verbeek, and Cordelia Schmid · 2013
Earlier work this paper cites.
Tv-l1 optical flow estimation
Javier Sánchez Pérez, Enric Meinhardt-Llopis, and Gabriele Facciolo · 2013
Earlier work this paper cites.
A survey on activity recognition and behavior understanding in video surveillance
Sarvesh Vishwakarma and Anupam Agrawal · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Ssd: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2016
Cited alongside, same era.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2016
Cited alongside, same era.
End-to-end, single-stream temporal action detection in untrimmed videos
Shyamal Buch, Victor Escorcia, Bernard Ghanem, Li Fei-Fei, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Rethinking the faster r-cnn architecture for temporal action localization
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A Ross, Jia Deng, and Rahul Sukthankar · 2018
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Diederik P Kingma and Prafulla Dhariwal · 2018
Later among the works it cites.
Recurrent tubelet proposal and recognition networks for action detection
Dong Li, Zhaofan Qiu, Qi Dai, Ting Yao, and Tao Mei · 2018
Later among the works it cites.
Bsn: Boundary sensitive network for temporal action proposal generation
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang · 2018
Later among the works it cites.
Weakly supervised action localization by sparse temporal pooling network
Phuc Nguyen, Ting Liu, Gautam Prasad, and Bohyung Han · 2018
Later among the works it cites.
W-talc: Weakly-supervised temporal activity localization and classification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Temporal context network for activity localization in videos
Xiyang Dai, Bharat Singh, Guyue Zhang, Larry S Davis, and Yan Qiu Chen · 2017
Cited alongside, same era.
Cascaded boundary regression for temporal action detection
Jiyang Gao, Zhenheng Yang, and Ram Nevatia · 2017
Cited alongside, same era.
Scc: Semantic context cascade for efficient action detection
Fabian Caba Heilbron, Wayner Barrios, Victor Escorcia, and Bernard Ghanem · 2017
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Cited alongside, same era.
The thumos challenge on action recognition for videos “in the wild”
Haroon Idrees, Amir R Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Single shot temporal action detection
Tianwei Lin, Xu Zhao, and Zheng Shou · 2017
Cited alongside, same era.
Sujoy Paul, Sourya Roy, and Amit K Roy-Chowdhury · 2018
Later among the works it cites.
Autoloc: Weakly-supervised temporal action localization in untrimmed videos
Zheng Shou, Hang Gao, Lei Zhang, Kazuyuki Miyazawa, and Shih-Fu Chang · 2018
Later among the works it cites.
S3d: Single shot multi-span detector via fully 3d convolutional network
Da Zhang, Xiyang Dai, Xin Wang, and Yuan-Fang Wang · 2018
Later among the works it cites.
Step-by-step erasion, one-by-one collection: A weakly supervised temporal action detector
Jia-Xing Zhong, Nannan Li, Weijie Kong, Tao Zhang, Thomas H Li, and Ge Li · 2018
Later among the works it cites.
Long short-term relation networks for video action detection
Dong Li, Ting Yao, Zhaofan Qiu, Houqiang Li, and Tao Mei · 2019
Later among the works it cites.
Completeness modeling and context separation for weakly supervised temporal action localization
Daochang Liu, Tingting Jiang, and Yizhou Wang · 2019
Later among the works it cites.
Weakly supervised temporal action localization through contrast based evaluation networks
Ziyi Liu, Le Wang, Qilin Zhang, Zhanning Gao, Zhenxing Niu, Nanning Zheng, and Gang Hua · 2019
Later among the works it cites.
Gaussian temporal awareness networks for action localization
Fuchen Long, Ting Yao, Zhaofan Qiu, Xinmei Tian, Jiebo Luo, and Tao Mei · 2019
Later among the works it cites.
3c-net: Category count and center loss for weakly-supervised action localization
Sanath Narayan, Hisham Cholakkal, Fahad Shahbaz Khan, and Ling Shao · 2019
Later among the works it cites.
Weakly-supervised action localization with background modeling
Phuc Xuan Nguyen, Deva Ramanan, and Charless C Fowlkes · 2019
Later among the works it cites.
Learning spatio-temporal representation with local and global diffusion
Zhaofan Qiu, Ting Yao, Chong-Wah Ngo, Xinmei Tian, and Tao Mei · 2019
Later among the works it cites.
Less is more: Learning highlight detection from video duration
Bo Xiong, Yannis Kalantidis, Deepti Ghadiyaram, and Kristen Grauman · 2019
Later among the works it cites.
Temporal structure mining for weakly supervised action detection
Tan Yu, Zhou Ren, Yuncheng Li, Enxu Yan, Ning Xu, and Junsong Yuan · 2019
Later among the works it cites.
Marginalized average attentional network for weakly-supervised learning
Yuan Yuan, Yueming Lyu, Xi Shen, Ivor W. Tsang, and Dit-Yan Yeung · 2019
Later among the works it cites.
Graph convolutional networks for temporal action localization
Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan · 2019
Later among the works it cites.