Fetching the paper…
Reading the bibliography…
Due to the large memory footprint of untrimmed videos, current state-of-the-art video localization methods operate atop precomputed video clip features.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Temporal localization of actions with actoms
Adrien Gaidon, Zaid Harchaoui, and Cordelia Schmid · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Efficient action localization with approximately normalized fisher vectors
Dan Oneata, Jakob Verbeek, and Cordelia Schmid · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Daps: Deep action proposals for action understanding
Victor Escorcia, Fabian Caba Heilbron, Juan Carlos Niebles, and Bernard Ghanem · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2016
Earlier work this paper cites.
Image captioning with semantic attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo · 2016
Earlier work this paper cites.
Sst: Single-stream temporal action proposals
Shyamal Buch, Victor Escorcia, Chuanqi Shen, Bernard Ghanem, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Scc: Semantic context cascade for efficient action detection
Fabian Caba Heilbron, Wayner Barrios, Victor Escorcia, and Bernard Ghanem · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Temporal context network for activity localization in videos
Xiyang Dai, Bharat Singh, Guyue Zhang, Larry S Davis, and Yan Qiu Chen · 2017
Earlier work this paper cites.
Temporal context network for activity localization in videos
Xiyang Dai, Bharat Singh, Guyue Zhang, Larry S Davis, and Yan Qiu Chen · 2017
Earlier work this paper cites.
Turn tap: Temporal unit regression network for temporal action proposals
Jiyang Gao, Zhenheng Yang, Kan Chen, Chen Sun, and Ram Nevatia · 2017
Earlier work this paper cites.
Cascaded boundary regression for temporal action detection
Jiyang Gao, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Real-time temporal action localization in untrimmed videos by sub-action discovery
Rui Hou, Rahul Sukthankar, and Mubarak Shah · 2017
Earlier work this paper cites.
The thumos challenge on action recognition for videos “in the wild”
Haroon Idrees, Amir R Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization
Krishna Kumar Singh and Yong Jae Lee · 2017
Earlier work this paper cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2017
Earlier work this paper cites.
Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos
Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang · 2017
Earlier work this paper cites.
End-to-end, single-stream temporal action detection in untrimmed videos
Bernard Ghanem Shyamal Buch, Victor Escorcia and Juan Carlos Niebles · 2017
Cited alongside, same era.
R-c3d: Region convolutional 3d network for temporal activity detection
Huijuan Xu, Abir Das, and Kate Saenko · 2017
Cited alongside, same era.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin · 2017
Cited alongside, same era.
Diagnosing error in temporal action detectors
Humam Alwassel, Fabian Caba Heilbron, Victor Escorcia, and Bernard Ghanem · 2018
Cited alongside, same era.
Action search: Spotting targets in videos and its application to temporal action localization
Humam Alwassel, Fabian Caba Heilbron, and Bernard Ghanem · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Weakly supervised temporal action localization through contrast based evaluation networks
Ziyi Liu, Le Wang, Qilin Zhang, Zhanning Gao, Zhenxing Niu, Nanning Zheng, and Gang Hua · 2019
Later among the works it cites.
Gaussian temporal awareness networks for action localization
Fuchen Long, Ting Yao, Zhaofan Qiu, Xinmei Tian, Jiebo Luo, and Tao Mei · 2019
Later among the works it cites.
Streamlined dense video captioning
Jonghwan Mun, Linjie Yang, Zhou Ren, Ning Xu, and Bohyung Han · 2019
Later among the works it cites.
Memory-attended recurrent network for video captioning
Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai · 2019
Later among the works it cites.
Watch, listen and tell: Multi-modal weakly supervised dense event captioning
Tanzila Rahman, Bicheng Xu, and Leonid Sigal · 2019
Later among the works it cites.
Dense procedure captioning in narrated instructional videos
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Contextual multi-scale region convolutional 3d network for activity detection
Yancheng Bai, Huijuan Xu, Kate Saenko, and Bernard Ghanem · 2018
Cited alongside, same era.
Rethinking the faster r-cnn architecture for temporal action localization
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A Ross, Jia Deng, and Rahul Sukthankar · 2018
Cited alongside, same era.
Encoder-decoder with atrous separable convolution for semantic image segmentation
Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam · 2018
Cited alongside, same era.
Ctap: Complementary temporal action proposal generation
Jiyang Gao, Kan Chen, and Ram Nevatia · 2018
Cited alongside, same era.
The activitynet large-scale activity recognition challenge 2018 summary
Bernard Ghanem, Juan Carlos Niebles, Cees Snoek, Fabian Caba Heilbron, Humam Alwassel, Victor Escorcia, Ranjay Khrisna, Shyamal Buch, and Cuong Duc Dao · 2018
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Botian Shi, Lei Ji, Yaobo Liang, Nan Duan, Peng Chen, Zhendong Niu, and Ming Zhou · 2019
Later among the works it cites.
Video classification with channel-separated convolutional networks
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Later among the works it cites.
Controllable video captioning with pos sequence guidance based on gated fusion network
Bairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang, Jingwen Wang, and Wei Liu · 2019
Later among the works it cites.
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, and Wei Liu · 2019
Later among the works it cites.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Later among the works it cites.
Text-to-clip video retrieval with early fusion and re-captioning
Huijuan Xu, Kun He, Leonid Sigal, Stan Sclaroff, and Kate Saenko · 2019
Later among the works it cites.
Stat: spatial-temporal attention mechanism for video captioning
Chenggang Yan, Yunbin Tu, Xingzheng Wang, Yongbing Zhang, Xinhong Hao, Yongdong Zhang, and Qionghai Dai · 2019
Later among the works it cites.
Graph convolutional networks for temporal action localization
Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan · 2019
Later among the works it cites.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Closest in time.
Boundary content graph neural network for temporal action proposal generation
Yueran Bai, Yingying Wang, Yunhai Tong, Yang Yang, Qiyue Liu, and Junhui Liu · 2020
Closest in time.
Accurate temporal action proposal generation with relation-aware pyramid network
Jialin Gao, Zhixiang Shi, Guanshuo Wang, Jiani Li, Yufeng Yuan, Shiming Ge, and Xi Zhou · 2020
Closest in time.
Scale matters: Temporal scale aggregation network for precise action localization in untrimmed videos
Guoqiang Gong, Liangfeng Zheng, and Yadong Mu · 2020
Closest in time.
A better use of audio-visual cues: Dense video captioning with bi-modal transformer
Vladimir Iashin and Esa Rahtu · 2020
Closest in time.
Multi-modal dense video captioning
Vladimir Iashin and Esa Rahtu · 2020
Closest in time.
Actionbytes: Learning from trimmed videos to localize actions
Mihir Jain, Amir Ghodrati, and Cees G. M. Snoek · 2020
Closest in time.
Background suppression network for weakly-supervised temporal action localization
Pilhyeon Lee, Youngjung Uh, and Hyeran Byun · 2020
Closest in time.
Deep concept-wise temporal convolutional networks for action localization
Xin Li, Tianwei Lin, Xiao Liu, Wangmeng Zuo, Chao Li, Xiang Long, Dongliang He, Fu Li, Shilei Wen, and Chuang Gan · 2020
Closest in time.
Fast learning of temporal action proposal via dense boundary generator
Chuming Lin, Jian Li, Yabiao Wang, Ying Tai, Donghao Luo, Zhipeng Cui, Chengjie Wang, Jilin Li, Feiyue Huang, and Rongrong Ji · 2020
Closest in time.
Progressive boundary refinement network for temporal action detection
Qinying Liu and Zilei Wang · 2020
Closest in time.
SF-Net: Single-Frame Supervision for Temporal Action Localization
Fan Ma, Linchao Zhu, Yi Yang, Shengxin Zha, Gourab Kundu, Matt Feiszli, and Zheng Shou · 2020
Closest in time.
Spatio-temporal graph for video captioning with knowledge distillation
Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, and Juan Carlos Niebles · 2020
Closest in time.
Efficientdet: Scalable and efficient object detection
Mingxing Tan, Ruoming Pang, and Quoc V Le · 2020
Closest in time.
G-tad: Sub-graph localization for temporal action detection
Mengmeng Xu, Chen Zhao, David S. Rojas, Ali Thabet, and Bernard Ghanem · 2020
Closest in time.
Bottom-up temporal action localization with mutual regularization
Peisen Zhao, Lingxi Xie, Chen Ju, Ya Zhang, Yanfeng Wang, and Qi Tian · 2020
Closest in time.
Syntax-aware action targeting for video captioning
Qi Zheng, Chaoyue Wang, and Dacheng Tao · 2020
Closest in time.
Refineloc: Iterative refinement for weakly-supervised action localization
Alejandro Pardo, Humam Alwassel, Fabian Caba Heilbron, Ali Thabet, and Bernard Ghanem · 2021
Closest in time.