Fetching the paper…
Reading the bibliography…
This paper tackles an emerging and challenging problem of long video temporal grounding~(VTG) that localizes video moments related to a natural language (NL) query.
Excl: Extractive clip localization using natural language descriptions
Soham Ghosh, Anuva Agarwal, Zarana Parekh, and Alexander G. Hauptmann. 2019 · 1990
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017 · 2017
Earlier work this paper cites.
TALL: temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017 · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Movie description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele. 2017 · 2017
Earlier work this paper cites.
TVQA: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg. 2018 · 2018
Earlier work this paper cites.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick. 2019 · 2019
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020 · 2020
Earlier work this paper cites.
Learning to contrast the counterfactual samples for robust visual question answering
Zujie Liang, Weitao Jiang, Haifeng Hu, and Jiaying Zhu. 2020 · 2020
Earlier work this paper cites.
Self-supervised learning of pretext-invariant representations
Ishan Misra and Laurens van der Maaten. 2020 · 2020
Earlier work this paper cites.
Dense regression network for video grounding
Runhao Zeng, Haoming Xu, Wenbing Huang, Peihao Chen, Mingkui Tan, and Chuang Gan. 2020 · 2020
Earlier work this paper cites.
On pursuit of designing multi-modal transformer for video grounding
Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, and Yuexian Zou. 2021 · 2021
Cited alongside, same era.
SimCSE: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021b · 2021
Cited alongside, same era.
Conquer: Contextual query-aware ranking for video corpus moment retrieval
Zhijian Hou, Chong-Wah Ngo, and Wing Kwong Chan. 2021 · 2021
Cited alongside, same era.
Detecting moments and highlights in videos via natural language queries
Jie Lei, Tamara L Berg, and Mohit Bansal. 2021 · 2021
Cited alongside, same era.
Clip4clip: An empirical study of clip for end to end video clip retrieval
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2021 · 2021
Cited alongside, same era.
Tallformer: Temporal action localization with a long-memory transformer
Gedas Bertasius Feng Cheng. 2022 · 2022
Closest in time.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. 2022 · 2022
Closest in time.
Long movie clip classification with state-space video models
Md Mohaiminul Islam and Gedas Bertasius. 2022 · 2022
Closest in time.
Mad: A scalable dataset for language grounding in videos from movie audio descriptions
Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba, Chen Zhao, Silvio Giancola, and Bernard Ghanem. 2022 · 2022
Closest in time.
Contrastive video-language learning with fine-grained frame sampling
Zixu Wang, Yujie Zhong, Yishu Miao, Lin Ma, and Lucia Specia. 2022 · 2022
Closest in time.
Tubedetr: Spatio-temporal video grounding with transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Coco-lm: Correcting and contrasting text sequences for language model pretraining
Yu Meng, Chenyan Xiong, Payal Bajaj, Paul Bennett, Jiawei Han, Xia Song, et al. 2021 · 2021
Cited alongside, same era.
Interventional video grounding with dual contrastive learning
Guoshun Nan, Rui Qiao, Yao Xiao, Jun Liu, Sicong Leng, Hao Zhang, and Wei Lu. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Cited alongside, same era.
Vlg-net: Video-language graph matching network for video grounding
Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner, and Bernard Ghanem. 2021 · 2021
Cited alongside, same era.
Towards long-form video understanding
Chao-Yuan Wu and Philipp Krahenbuhl. 2021 · 2021
Cited alongside, same era.
VideoCLIP: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021 · 2021
Cited alongside, same era.
Embracing uncertainty: Decoupling and de-bias for robust temporal grounding
Hao Zhou, Chongyang Zhang, Yan Luo, Yanjun Chen, and Chuanping Hu. 2021 · 2021
Cited alongside, same era.
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. 2022 · 2022
Closest in time.
The elements of temporal sentence grounding in videos: A survey and future directions
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. 2022 · 2022
Closest in time.
Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learning
Minghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng, and Yang Liu. 2022 · 2022
Closest in time.
Liveseg: Unsupervised multimodal temporal segmentation of long livestream videos
Jielin Qiu, Franck Dernoncourt, Trung Bui, Zhaowen Wang, Ding Zhao, and Hailin Jin. 2023 · 2023
Closest in time.
Naq: Leveraging narrations as queries to supervise episodic memory
Santhosh Kumar Ramakrishnan, Ziad Al-Halah, and Kristen Grauman. 2023 · 2023
Closest in time.
Control image captioning spatially and temporally
Kun Yan, Lei Ji, Huaishao Luo, Ming Zhou, Nan Duan, and Shuai Ma. 2021 · 2025
Closest in time.