Fetching the paper…
Reading the bibliography…
Video activity localization aims at understanding the semantic content in long untrimmed videos and retrieving actions of interest.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
Grounding Action Descriptions in Videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Thumos challenge: Action recognition with a large number of classes, 2014
Yu-Gang Jiang, Jingen Liu, A Roshan Zamir, George Toderici, Ivan Laptev, Mubarak Shah, and Rahul Sukthankar · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Localizing Moments in Video With Natural Language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
TALL: Temporal Activity Localization via Language Query
Gao Jiyang, Sun Chen, Yang Zhenheng, Nevatia, Ram · 2017
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
CDC: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos
Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang · 2017
Earlier work this paper cites.
Deep learning for video classification and captioning
Zuxuan Wu, Ting Yao, Yanwei Fu, and Yu-Gang Jiang · 2017
Earlier work this paper cites.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin · 2017
Earlier work this paper cites.
Diagnosing error in temporal action detectors
Humam Alwassel, Fabian Caba Heilbron, Victor Escorcia, and Bernard Ghanem · 2018
Earlier work this paper cites.
Temporally Grounding Natural Sentence in Video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Bsn: Boundary sensitive network for temporal action proposal generation
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang · 2018
Earlier work this paper cites.
Attentive Moment Retrieval in Videos
Meng Liu, Xiang Wang, Liqiang Nie, Xiangnan He, Baoquan Chen, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Cross-Modal Moment Localization in Videos
Meng Liu, Xiang Wang, Liqiang Nie, Qi Tian, Baoquan Chen, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Deepdecision: A mobile deep learning framework for edge video analytics
Xukan Ran, Haolianz Chen, Xiaodan Zhu, Zhenming Liu, and Jiasi Chen · 2018
Earlier work this paper cites.
VAL: Visual-Attention Action Localizer
Xiaomeng Song and Yahong Han · 2018
Earlier work this paper cites.
Multi-modal Circulant Fusion for Video-to-Language and Backward
Aming Wu and Yahong Han · 2018
Earlier work this paper cites.
Semantic Proposal for Activity Localization in Videos via Sentence Query
Shaoxiang Chen and Yu-Gang Jiang · 2019
Earlier work this paper cites.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan C. Russell · 2019
Earlier work this paper cites.
MAC: Mining Activity Concepts for Language-based Temporal Localization
Runzhou Ge, Jiyang Gao, Kan Chen, and Ram Nevatia · 2019
Earlier work this paper cites.
Read, watch, and move: Reinforcement learning for temporally grounding natural language descriptions in videos
Dongliang He, Xiang Zhao, Jizhou Huang, Fu Li, Xiao Liu, and Shilei Wen · 2019
Cited alongside, same era.
Cross-Modal Video Moment Retrieval with Spatial and Language-Temporal Attention
Bin Jiang, Xin Huang, Chao Yang, and Junsong Yuan · 2019
Cited alongside, same era.
BMN: boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Generalized intersection over union: A metric and a loss for bounding box regression
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese · 2019
Cited alongside, same era.
Learning salient boundary feature for anchor-free temporal action localization
Chuming Lin, Chengming Xu, Donghao Luo, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yanwei Fu · 2021
Later among the works it cites.
Interventional video grounding with dual contrastive learning
Guoshun Nan, Rui Qiao, Yao Xiao, Jun Liu, Sicong Leng, Hao Zhang, and Wei Lu · 2021
Later among the works it cites.
Video transformer network
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann · 2021
Later among the works it cites.
Vlg-net: Video-language graph matching network for video grounding
Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner, and Bernard Ghanem · 2021
Later among the works it cites.
Sparse r-cnn: End-to-end object detection with learnable proposals
Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, et al · 2021
Later among the works it cites.
Relaxed transformer decoders for direct action proposal generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weining Wang, Yan Huang, and Liang Wang · 2019
Cited alongside, same era.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Hierarchical Visual-Textual Graph for Temporal Activity Localization via Language
Chen Shaoxiang, Jiang Yu-Gang · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Cited alongside, same era.
Moment Retrieval via Cross-Modal Interaction Networks With Query Reconstruction
Zhijie Lin, Zhou Zhao, Zhu Zhang, Zijian Zhang, and Deng Cai · 2020
Cited alongside, same era.
Jing Tan, Jiaqi Tang, Limin Wang, and Gangshan Wu · 2021
Later among the works it cites.
Rgb stream is enough for temporal action detection
Chenhao Wang, Hongxiang Cai, Yuxin Zou, and Yichao Xiong · 2021
Later among the works it cites.
Uncertainty guided collaborative training for weakly supervised temporal action detection
Wenfei Yang, Tianzhu Zhang, Xiaoyuan Yu, Tian Qi, Yongdong Zhang, and Feng Wu · 2021
Later among the works it cites.
Multi-stage aggregated transformer network for temporal language localization in videos
Mingxing Zhang, Yang Yang, Xinghan Chen, Yanli Ji, Xing Xu, Jingjing Li, and Heng Tao Shen · 2021
Later among the works it cites.
Video self-stitching graph network for temporal action localization
Chen Zhao, Ali K Thabet, and Bernard Ghanem · 2021
Later among the works it cites.
Internvideo-ego4d: A pack of champion solutions to ego4d challenges, 2022
Guo Chen, Sen Xing, Zhe Chen, Yi Wang, Kunchang Li, Yizhuo Li, Yi Liu, Jiahao Wang, Yin-Dong Zheng, Bingkun Huang, Zhiyu Zhao, Junting Pan, Yifei Huang, Zun Wang, Jiashuo Yu, Yinan He, Hongjie Zhang, Tong Lu, Yali Wang, Limin Wang, and Yu Qiao · 2022
Later among the works it cites.
Diffusiondet: Diffusion model for object detection
Shoufa Chen, Peize Sun, Yibing Song, and Ping Luo · 2022
Later among the works it cites.
Tallformer: Temporal action localization with a long-memory transformer
Feng Cheng and Gedas Bertasius · 2022
Later among the works it cites.
Cone: An efficient coarse-to-fine alignment framework for long video temporal grounding
Zhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, Wing-Kwong Chan, Chong-Wah Ngo, Zheng Shou, and Nan Duan · 2022
Later among the works it cites.
Video activity localisation with uncertainties in temporal boundary
Jiabo Huang, Hailin Jin, Shaogang Gong, and Yang Liu · 2022
Later among the works it cites.
Dn-detr: Accelerate detr training by introducing query denoising
Feng Li, Hao Zhang, Shilong Liu, Jian Guo, Lionel M Ni, and Lei Zhang · 2022
Later among the works it cites.
Exploring denoised cross-video contrast for weakly-supervised temporal action localization
Jingjing Li, Tianyu Yang, Wei Ji, Jue Wang, and Li Cheng · 2022
Later among the works it cites.
Egocentric video-language pretraining
Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Later among the works it cites.
Reler@zju-alibaba submission to the ego4d natural language queries challenge 2022, 2022
Naiyuan Liu, Xiaohan Wang, Xiaobo Li, Yi Yang, and Yueting Zhuang · 2022
Later among the works it cites.
An empirical study of end-to-end temporal action detection
Xiaolong Liu, Song Bai, and Xiang Bai · 2022
Later among the works it cites.
Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection
Ye Liu, Siyuan Li, Yang Wu, Chang-Wen Chen, Ying Shan, and Xiaohu Qie · 2022
Later among the works it cites.
Mad: A scalable dataset for language grounding in videos from movie audio descriptions
Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba, Chen Zhao, Silvio Giancola, and Bernard Ghanem · 2022
Later among the works it cites.
Uncertainty guided collaborative training for weakly supervised and unsupervised temporal action localization
Wenfei Yang, Tianzhu Zhang, Yongdong Zhang, and Feng Wu · 2022
Later among the works it cites.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel Ni, and Heung-Yeung Shum · 2022
Later among the works it cites.
Localizing moments in long video via multimodal guidance
Wayner Barrios, Mattia Soldan, Fabian Caba Heilbron, Alberto Mario Ceballos-Arroyo, and Bernard Ghanem · 2023
Closest in time.