Fetching the paper…
Reading the bibliography…
Video Temporal Grounding (VTG), which aims to ground target clips from videos (such as consecutive intervals or disjoint shots) according to custom language queries (e.g., sentences or words), is key for video browsing on social media.
Automatically extracting highlights for tv baseball programs
Yong Rui, Anoop Gupta, and Alex Acero · 2000
Earlier work this paper cites.
Discovering important people and objects for egocentric video summarization
Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman · 2012
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Creating summaries from user videos
Michael Gygli, Helmut Grabner, Hayko Riemenschneider, and Luc Van Gool · 2014
Earlier work this paper cites.
Category-specific video summarization
Danila Potapov, Matthijs Douze, Zaid Harchaoui, and Cordelia Schmid · 2014
Earlier work this paper cites.
Ranking domain-specific highlights by analyzing edited videos
Min Sun, Ali Farhadi, and Steve Seitz · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Multi-task deep visual-semantic embedding for video thumbnail selection
Wu Liu, Tao Mei, Yongdong Zhang, Cherry Che, and Jiebo Luo · 2015
Earlier work this paper cites.
Tvsum: Summarizing web videos using titles
Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes · 2015
Earlier work this paper cites.
Unsupervised extraction of video highlights via robust recurrent auto-encoders
Huan Yang, Baoyuan Wang, Stephen Lin, David Wipf, Minyi Guo, and Baining Guo · 2015
Earlier work this paper cites.
Video2gif: Automatic generation of animated gifs from video
Michael Gygli, Yale Song, and Liangliang Cao · 2016
Earlier work this paper cites.
To click or not to click: Automatic selection of beautiful thumbnails from videos
Yale Song, Miriam Redi, Jordi Vallmitjana, and Alejandro Jaimes · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Video summarization with long short-term memory
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Unsupervised video summarization with adversarial lstm networks
Behrooz Mahasseni, Michael Lam, and Sinisa Todorovic · 2017
Earlier work this paper cites.
Query-focused video summarization: Dataset, evaluation, and a memory network based approach
Aidean Sharghi, Jacob S Laurel, and Boqing Gong · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
Excl: Extractive clip localization using natural language descriptions
Soham Ghosh, Anuva Agarwal, Zarana Parekh, and Alexander G. Hauptmann · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Generalized intersection over union: A metric and a loss for bounding box regression
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese · 2019
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Boundary-sensitive pre-training for temporal localization in videos
Mengmeng Xu, Juan-Manuel Pérez-Rúa, Victor Escorcia, Brais Martinez, Xiatian Zhu, Li Zhang, Bernard Ghanem, and Tao Xiang · 2021
Later among the works it cites.
Cross-category video highlight detection via set-based learning
Minghao Xu, Hang Wang, Bingbing Ni, Riheng Zhu, Zhenbang Sun, and Changhu Wang · 2021
Later among the works it cites.
Temporal cue guided video highlight detection with low-rank audio-visual fusion
Qinghao Ye, Xiyue Shen, Yuan Gao, Zirui Wang, Qi Bi, Ping Li, and Guang Yang · 2021
Later among the works it cites.
Locvtp: Video-text pre-training for temporal localization
Meng Cao, Tianyu Yang, Junwu Weng, Can Zhang, Jue Wang, and Yuexian Zou · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Objects365: A large-scale, high-quality dataset for object detection
Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun · 2019
Cited alongside, same era.
Less is more: Learning highlight detection from video duration
Bo Xiong, Yannis Kalantidis, Deepti Ghadiyaram, and Kristen Grauman · 2019
Cited alongside, same era.
Mini-net: Multiple instance ranking network for video highlight detection
Fa-Ting Hong, Xuanteng Huang, Wei-Hong Li, and Wei-Shi Zheng · 2020
Cited alongside, same era.
Tvr: A large-scale dataset for video-subtitle moment retrieval
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Cited alongside, same era.
Local-global video-text interactions for temporal grounding
Jonghwan Mun, Minsu Cho, and Bohyung Han · 2020
Cited alongside, same era.
Watch hours in minutes: Summarizing videos with user intent
Saiteja Nalla, Mohit Agrawal, Vishal Kaushal, Ganesh Ramakrishnan, and Rishabh Iyer · 2020
Cited alongside, same era.
Learning trailer moments in full-length movies with co-contrastive attention
Lezi Wang, Dong Liu, Rohit Puri, and Dimitris N Metaxas · 2020
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
Joint video summarization and moment localization by cross-task sample transfer
Hao Jiang and Yadong Mu · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
Grounded language-image pre-training
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al · 2022
Later among the works it cites.
Egocentric video-language pretraining
Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Later among the works it cites.
Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection
Ye Liu, Siyuan Li, Yang Wu, Chang-Wen Chen, Ying Shan, and Xiaohu Qie · 2022
Later among the works it cites.
Learning audio-video modalities from image captions
Arsha Nagrani, Paul Hongsuck Seo, Bryan Seybold, Anja Hauth, Santiago Manen, Chen Sun, and Cordelia Schmid · 2022
Later among the works it cites.
Volta: Vision-language transformer with weakly-supervised local-feature alignment
Shraman Pramanick, Li Jing, Sayan Nag, Jiachen Zhu, Hardik Shah, Yann LeCun, and Rama Chellappa · 2022
Later among the works it cites.
Intentvizor: Towards generic query guided interactive video summarization
Guande Wu, Jianzhe Lin, and Claudio T Silva · 2022
Later among the works it cites.
Unsupervised pre-training for temporal action localization tasks
Can Zhang, Tianyu Yang, Junwu Weng, Meng Cao, Jue Wang, and Yuexian Zou · 2022
Later among the works it cites.
Glipv2: Unifying localization and vision-language understanding
Haotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen, Liunian Harold Li, Xiyang Dai, Lijuan Wang, Lu Yuan, Jenq-Neng Hwang, and Jianfeng Gao · 2022
Later among the works it cites.
Query-dependent video representation for moment retrieval and highlight detection
WonJun Moon, Sangeek Hyun, SangUk Park, Dongchan Park, and Jae-Pil Heo · 2023
Closest in time.
Egovlpv2: Egocentric video-language pre-training with fusion in the backbone
Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang · 2023
Closest in time.
Naq: Leveraging narrations as queries to supervise episodic memory
Santhosh Kumar Ramakrishnan, Ziad Al-Halah, and Kristen Grauman · 2023
Closest in time.
All in one: Exploring unified video-language pre-training
Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Kevin Qinghong Lin, Satoshi Tsutsui, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, et al · 2023
Closest in time.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning, 2023
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid · 2023
Closest in time.
Learning video representations from large language models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar · 2023
Closest in time.