Fetching the paper…
Reading the bibliography…
The query-based moment retrieval is a problem of localising a specific clip from an untrimmed video according a query sentence.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
TALL: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Cited alongside, same era.
What actions are needed for understanding human actions in videos?
Gunnar A. Sigurdsson, Olga Russakovsky, and Abhinav Gupta · 2017
Cited alongside, same era.
Diagnosing error in temporal action detectors
Humam Alwassel, Fabian Caba Heilbron, Victor Escorcia, and Bernard Ghanem · 2018
Cited alongside, same era.
From same photo: Cheating on visual kinship challenges
Mitchell Dawson, Andrew Zisserman, and Christoffer Nellåker · 2018
Cited alongside, same era.
Attentive moment retrieval in videos
Meng Liu, Xiang Wang, Liqiang Nie, Xiangnan He, Baoquan Chen, and Tat-Seng Chua · 2018
Cited alongside, same era.
Semantic proposal for activity localization in videos via sentence query
Shaoxiang Chen and Yu-Gang Jiang · 2019
Cited alongside, same era.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell · 2019
Later among the works it cites.
Tripping through time: Efficient localization of activities in videos
Meera Hahn, Asim Kadav, James M. Rehg, and Hans Peter Graf · 2019
Later among the works it cites.
Cross-modal video moment retrieval with spatial and language-temporal attention
Bin Jiang, Xin Huang, Chao Yang, and Junsong Yuan · 2019
Later among the works it cites.
Language-driven temporal activity localization: A semantic matching reinforcement learning model
Weining Wang, Yan Huang, and Liang Wang · 2019
Later among the works it cites.
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan A. Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semantic conditioned dynamic modulation for temporal sentence grounding in videos
Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu
Cited in the paper.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yitian Yuan, Tao Mei, and Wenwu Zhu
Cited in the paper.
Later among the works it cites.
Learning 2D temporal adjacent networks for moment localization with natural language
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo · 2020
Closest in time.