Fetching the paper…
Reading the bibliography…
Given an untrimmed video and a natural language query, Natural Language Video Localization (NLVL) aims to identify the video moment described by the query.
Excl: Extractive clip localization using natural language descriptions
Soham Ghosh, Anuva Agarwal, Zarana Parekh, and Alexander G. Hauptmann. 2019 · 1990
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir D. Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Earlier work this paper cites.
R-FCN: object detection via region-based fully convolutional networks
Jifeng Dai, Yi Li, Kaiming He, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
João Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
TALL: temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017 · 2017
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell. 2017 · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Video question answering via attribute-augmented attention network learning
Yunan Ye, Zhou Zhao, Yimeng Li, Long Chen, Jun Xiao, and Yueting Zhuang. 2017 · 2017
Earlier work this paper cites.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua. 2018 · 2018
Earlier work this paper cites.
Localizing moments in video with temporal language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell. 2018 · 2018
Earlier work this paper cites.
Cornernet: Detecting objects as paired keypoints
Hei Law and Jia Deng. 2018 · 2018
Earlier work this paper cites.
TVQA: localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L. Berg. 2018 · 2018
Cited alongside, same era.
Find and focus: Retrieve and localize video events with natural language queries
Dian Shao, Yu Xiong, Yue Zhao, Qingqiu Huang, Yu Qiao, and Dahua Lin. 2018 · 2018
Cited alongside, same era.
Qanet: Combining local convolution with global self-attention for reading comprehension
Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, and Quoc V. Le. 2018 · 2018
Cited alongside, same era.
Semantic proposal for activity localization in videos via sentence query
Shaoxiang Chen and Yu-Gang Jiang. 2019 · 2019
Cited alongside, same era.
Centernet: Keypoint triplets for object detection
Kaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian. 2019 · 2019
Cited alongside, same era.
MAC: mining activity concepts for language-based temporal localization
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Later among the works it cites.
Rethinking the bottom-up framework for query-based video localization
Long Chen, Chujie Lu, Siliang Tang, Jun Xiao, Dong Zhang, Chilie Tan, and Xiaolin Li. 2020 · 2020
Later among the works it cites.
Local-global video-text interactions for temporal grounding
Jonghwan Mun, Minsu Cho, and Bohyung Han. 2020 · 2020
Later among the works it cites.
Proposal-free temporal moment localization of a natural-language query in video using guided attention
Cristian Rodriguez Opazo, Edison Marrese-Taylor, Fatemeh Sadat Saleh, Hongdong Li, and Stephen Gould. 2020 · 2020
Later among the works it cites.
Dense regression network for video grounding
Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Runzhou Ge, Jiyang Gao, Kan Chen, and Ram Nevatia. 2019 · 2019
Cited alongside, same era.
Read, watch, and move: Reinforcement learning for temporally grounding natural language descriptions in videos
Dongliang He, Xiang Zhao, Jizhou Huang, Fu Li, Xiao Liu, and Shilei Wen. 2019 · 2019
Cited alongside, same era.
DEBUG: A dense bottom-up grounding approach for natural language video localization
Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, and Jun Xiao. 2019 · 2019
Cited alongside, same era.
Language-driven temporal activity localization: A semantic matching reinforcement learning model
Weining Wang, Yan Huang, and Liang Wang. 2019 · 2019
Cited alongside, same era.
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan A. Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko. 2019 · 2019
Cited alongside, same era.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yitian Yuan, Tao Mei, and Wenwu Zhu. 2019 · 2019
Cited alongside, same era.
MAN: moment alignment network for natural language moment retrieval via iterative graph adjustment
Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S. Davis. 2019 · 2019
Cited alongside, same era.
Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, and Yuexian Zou. 2021 · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Closest in time.
Video relation detection via tracklet based visual transformer
Kaifeng Gao, Long Chen, Yifeng Huang, and Jun Xiao. 2021 · 2021
Closest in time.
Sparse R-CNN: end-to-end object detection with learnable proposals
Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, and Ping Luo. 2021 · 2021
Closest in time.
Crossformer: A versatile vision transformer based on cross-scale attention
Wenxiao Wang, Lu Yao, Long Chen, Deng Cai, Xiaofei He, and Wei Liu. 2021 · 2021
Closest in time.
Boundary proposal network for two-stage natural language video localization
Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, and Jun Xiao. 2021 · 2021
Closest in time.
A closer look at temporal sentence grounding in videos: Datasets and metrics
Yitian Yuan, Xiaohan Lan, Long Chen, Wei Liu, Xin Wang, and Wenwu Zhu. 2021 · 2021
Closest in time.