Fetching the paper…
Reading the bibliography…
The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques.
Video google: A text retrieval approach to object matching in videos
Josef Sivic and Andrew Zisserman · 2003
Earlier work this paper cites.
Memory dysfunction
Andrew E Budson and Bruce H Price · 2005
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper · 2009
Earlier work this paper cites.
The mediamill trecvid 2009 semantic video search engine
Cees Snoek, Kvd Sande, OD Rooij, Bouke Huurnink, J Uijlings, M van Liempt, M Bugalhoy, I Trancosoy, F Yan, M Tahir, et al · 2009
Earlier work this paper cites.
Grounding Action Descriptions in Videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele · 2015
Earlier work this paper cites.
Using descriptive video services to create a large data source for video annotation research
Atousa Torabi, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Earlier work this paper cites.
Towards episodic memory support for dementia patients by recognizing objects, faces and text in eye gaze
Takumi Toyama and Daniel Sonntag · 2015
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Bidirectional recurrent neural network with attention mechanism for punctuation restoration
Ottokar Tilk and Tanel Alumäe · 2016
Earlier work this paper cites.
Localizing Moments in Video With Natural Language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
TALL: Temporal Activity Localization via Language Query
Gao Jiyang, Sun Chen, Yang Zhenheng, Nevatia, Ram · 2017
Cited alongside, same era.
Dense-Captioning Events in Videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Movie description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele · 2017
Cited alongside, same era.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell · 2019
Cited alongside, same era.
Deepgcns: Can gcns go as deep as cnns?
Guohao Li, Matthias Müller, Ali Thabet, and Bernard Ghanem · 2019
Cited alongside, same era.
Deepergcn: All you need to train deeper gcns, 2020
Guohao Li, Chenxin Xiong, Ali Thabet, and Bernard Ghanem · 2020
Context-aware biaffine localizing network for temporal sentence grounding
Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou, Yu Cheng, Wei Wei, Zichuan Xu, and Yulai Xie · 2021
Closest in time.
Clip4clip: An empirical study of clip for end to end video clip retrieval
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2021
Closest in time.
Queryd: A video dataset with high-quality text and audio narrations
Andreea-Maria Oncescu, João F. Henriques, Yang Liu, Andrew Zisserman, and Samuel Albanie · 2021
Closest in time.
Learning to cut by watching movies
Alejandro Pardo, Fabian Caba, Juan León Alcázar, Ali K Thabet, and Bernard Ghanem · 2021
Closest in time.
Moviecuts: A new dataset and benchmark for cut type recognition
Alejandro Pardo, Fabian Caba Heilbron, Juan León Alcázar, Ali Thabet, and Bernard Ghanem · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Jointly cross-and self-modal graph attention network for query-based moment localization
Daizong Liu, Xiaoye Qu, Xiao-Yang Liu, Jianfeng Dong, Pan Zhou, and Zichuan Xu · 2020
Cited alongside, same era.
Local-Global Video-Text Interactions for Temporal Grounding
Jonghwan Mun, Minsu Cho, and Bohyung Han · 2020
Cited alongside, same era.
Uncovering hidden challenges in query-based video moment retrieval
Mayu Otani, Yuta Nakashima, Esa Rahtu, and Janne Heikkilä · 2020
Cited alongside, same era.
G-tad: Sub-graph localization for temporal action detection
Mengmeng Xu, Chen Zhao, David S. Rojas, Ali Thabet, and Bernard Ghanem · 2020
Cited alongside, same era.
Dense Regression Network for Video Grounding
Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan · 2020
Cited alongside, same era.
Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language
Zhang Songyang, Peng Houwen, Fu Jianlong, Luo, Jiebo · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
Vlg-net: Video-language graph matching network for video grounding
Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner, and Bernard Ghanem · 2021
Closest in time.
Transcript to video: Efficient clip sequencing from texts
Yu Xiong, Fabian Caba Heilbron, and Dahua Lin · 2021
Closest in time.
A closer look at temporal sentence grounding in videos: Datasets and metrics
Yitian Yuan, Xiaohan Lan, Long Chen, Wei Liu, Xin Wang, and Wenwu Zhu · 2021
Closest in time.
A closer look at temporal sentence grounding in videos: Datasets and metrics
Yitian Yuan, Xiaohan Lan, Long Chen, Wei Liu, Xin Wang, and Wenwu Zhu · 2021
Closest in time.
Towards debiasing temporal sentence grounding in video
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou · 2021
Closest in time.
Video self-stitching graph network for temporal action localization
Chen Zhao, Ali K Thabet, and Bernard Ghanem · 2021
Closest in time.
Cascaded prediction network via segment tree for temporal video grounding
Yang Zhao, Zhou Zhao, Zhu Zhang, and Zhijie Lin · 2021
Closest in time.
Embracing uncertainty: Decoupling and de-bias for robust temporal grounding
Hao Zhou, Chongyang Zhang, Yan Luo, Yanjun Chen, and Chuanping Hu · 2021
Closest in time.