Fetching the paper…
Reading the bibliography…
We aim to address the problem of Natural Language Video Localization (NLVL)-localizing the video segment corresponding to a natural language description in a long and untrimmed video.
ExCL: Extractive Clip Localization Using Natural Language Descriptions
Ghosh, S.; Agarwal, A.; Parekh, Z.; and Hauptmann, A. G. 2019 · 1990
Earlier work this paper cites.
Glove: Global Vectors for Word Representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R. B.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Learning Spatiotemporal Features with 3D Convolutional Networks
Tran, D.; Bourdev, L. D.; Fergus, R.; Torresani, L.; and Paluri, M. 2015 · 2015
Earlier work this paper cites.
R-FCN: Object Detection via Region-based Fully Convolutional Networks
Dai, J.; Li, Y.; He, K.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
TALL: Temporal Activity Localization via Language Query
Gao, J.; Sun, C.; Yang, Z.; and Nevatia, R. 2017 · 2017
Earlier work this paper cites.
Localizing Moments in Video with Natural Language
Hendricks, L. A.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. C. 2017 · 2017
Earlier work this paper cites.
Dense-Captioning Events in Videos
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Niebles, J. C. 2017 · 2017
Earlier work this paper cites.
Video question answering via attribute-augmented attention network learning
Ye, Y.; Zhao, Z.; Li, Y.; Chen, L.; Xiao, J.; and Zhuang, Y. 2017 · 2017
Earlier work this paper cites.
Temporally Grounding Natural Sentence in Video
Chen, J.; Chen, X.; Ma, L.; Jie, Z.; and Chua, T. 2018 · 2018
Earlier work this paper cites.
Localizing Moments in Video with Temporal Language
Hendricks, L. A.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. C. 2018 · 2018
Earlier work this paper cites.
CornerNet: Detecting Objects as Paired Keypoints
Law, H.; and Deng, J. 2018 · 2018
Cited alongside, same era.
TVQA: Localized, Compositional Video Question Answering
Lei, J.; Yu, L.; Bansal, M.; and Berg, T. L. 2018 · 2018
Cited alongside, same era.
Find and Focus: Retrieve and Localize Video Events with Natural Language Queries
Shao, D.; Xiong, Y.; Zhao, Y.; Huang, Q.; Qiao, Y.; and Lin, D. 2018 · 2018
Cited alongside, same era.
Multi-modal Circulant Fusion for Video-to-Language and Backward
Wu, A.; and Han, Y. 2018 · 2018
Cited alongside, same era.
QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension
Yu, A. W.; Dohan, D.; Luong, M.; Zhao, R.; Chen, K.; Norouzi, M.; and Le, Q. V. 2018 · 2018
Cited alongside, same era.
Localizing Natural Language in Videos
Chen, J.; Ma, L.; Chen, X.; Jie, Z.; and Luo, J. 2019 · 2019
Multilevel Language and Vision Integration for Text-to-Clip Retrieval
Xu, H.; He, K.; Plummer, B. A.; Sigal, L.; Sclaroff, S.; and Saenko, K. 2019 · 2019
Later among the works it cites.
To Find Where You Talk: Temporal Sentence Localization in Video with Attention Based Location Regression
Yuan, Y.; Mei, T.; and Zhu, W. 2019 · 2019
Later among the works it cites.
MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment
Zhang, D.; Dai, X.; Wang, X.; Wang, Y.; and Davis, L. S. 2019 · 2019
Later among the works it cites.
Rethinking the Bottom-Up Framework for Query-Based Video Localization
Chen, L.; Lu, C.; Tang, S.; Xiao, J.; Zhang, D.; Tan, C.; and Li, X. 2020 · 2020
Later among the works it cites.
Corner Proposal Network for Anchor-free, Two-stage Object Detection
Duan, K.; Xie, L.; Qi, H.; Bai, S.; Huang, Q.; and Tian, Q. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Semantic Proposal for Activity Localization in Videos via Sentence Query
Chen, S.; and Jiang, Y. 2019 · 2019
Cited alongside, same era.
CenterNet: Keypoint Triplets for Object Detection
Duan, K.; Bai, S.; Xie, L.; Qi, H.; Huang, Q.; and Tian, Q. 2019 · 2019
Cited alongside, same era.
MAC: Mining Activity Concepts for Language-Based Temporal Localization
Ge, R.; Gao, J.; Chen, K.; and Nevatia, R. 2019 · 2019
Cited alongside, same era.
Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
He, D.; Zhao, X.; Huang, J.; Li, F.; Liu, X.; and Wen, S. 2019 · 2019
Cited alongside, same era.
DEBUG: A Dense Bottom-Up Grounding Approach for Natural Language Video Localization
Lu, C.; Chen, L.; Tan, C.; Li, X.; and Xiao, J. 2019 · 2019
Cited alongside, same era.
Language-Driven Temporal Activity Localization: A Semantic Matching Reinforcement Learning Model
Wang, W.; Huang, Y.; and Wang, L. 2019 · 2019
Cited alongside, same era.
Mun, J.; Cho, M.; and Han, B. 2020 · 2020
Later among the works it cites.
Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention
Opazo, C. R.; Marrese-Taylor, E.; Saleh, F. S.; Li, H.; and Gould, S. 2020 · 2020
Later among the works it cites.
Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware Prediction
Wang, J.; Ma, L.; and Jiang, W. 2020 · 2020
Later among the works it cites.
Hierarchical Temporal Fusion of Multi-grained Attention Features for Video Question Answering
Xiao, S.; Li, Y.; Ye, Y.; Chen, L.; Pu, S.; Zhao, Z.; Shao, J.; and Xiao, J. 2020 · 2020
Later among the works it cites.
Dense Regression Network for Video Grounding
Zeng, R.; Xu, H.; Huang, W.; Chen, P.; Tan, M.; and Gan, C. 2020 · 2020
Later among the works it cites.
Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding
Chen, L.; Ma, W.; Xiao, J.; Zhang, H.; and Chang, S.-F. 2021 · 2021
Closest in time.