Fetching the paper…
Reading the bibliography…
Temporal Grounding is to identify specific moments or highlights from a video corresponding to textual descriptions.
Script data for attribute-based recognition of composite activities
Marcus Rohrbach, Michaela Regneri, Mykhaylo Andriluka, Sikandar Amin, Manfred Pinkal, and Bernt Schiele · 2012
Earlier work this paper cites.
Large-scale video summarization using web-image priors
Aditya Khosla, Raffay Hamid, Chih-Jen Lin, and Neel Sundaresan · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Ranking domain-specific highlights by analyzing edited videos
Min Sun, Ali Farhadi, and Steve Seitz · 2014
Earlier work this paper cites.
Multi-task deep visual-semantic embedding for video thumbnail selection
Wu Liu, Tao Mei, Yongdong Zhang, Cherry Che, and Jiebo Luo · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Tvsum: Summarizing web videos using titles
Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes · 2015
Earlier work this paper cites.
Unsupervised extraction of video highlights via robust recurrent auto-encoders
Huan Yang, Baoyuan Wang, Stephen Lin, David Wipf, Minyi Guo, and Baining Guo · 2015
Earlier work this paper cites.
To click or not to click: Automatic selection of beautiful thumbnails from videos
Yale Song, Miriam Redi, Jordi Vallmitjana, and Alejandro Jaimes · 2016
Earlier work this paper cites.
Video summarization with long short-term memory
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Learning visual n-grams from web data
Ang Li, Allan Jabri, Armand Joulin, and Laurens Van Der Maaten · 2017
Earlier work this paper cites.
Unsupervised video summarization with adversarial lstm networks
Behrooz Mahasseni, Michael Lam, and Sinisa Todorovic · 2017
Earlier work this paper cites.
Weakly supervised summarization of web videos
Rameswar Panda, Abir Das, Ziyan Wu, Jan Ernst, and Amit K Roy-Chowdhury · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Weakly-supervised video summarization using variational encoder-decoder and web prior
Sijia Cai, Wangmeng Zuo, Larry S Davis, and Lei Zhang · 2018
Earlier work this paper cites.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua · 2018
Earlier work this paper cites.
Localizing moments in video with temporal language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2018
Earlier work this paper cites.
Video summarization using fully convolutional sequence networks
Mrigank Rochan, Linwei Ye, and Yang Wang · 2018
Earlier work this paper cites.
Find and focus: Retrieve and localize video events with natural language queries
Dian Shao, Yu Xiong, Yue Zhao, Qingqiu Huang, Yu Qiao, and Dahua Lin · 2018
Cited alongside, same era.
Semantic proposal for activity localization in videos via sentence query
Shaoxiang Chen and Yu-Gang Jiang · 2019
Cited alongside, same era.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
Mac: Mining activity concepts for language-based temporal localization
Runzhou Ge, Jiyang Gao, Kan Chen, and Ram Nevatia · 2019
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Vlg-net: Video-language graph matching network for video grounding
Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner, and Bernard Ghanem · 2021
Later among the works it cites.
Image-text alignment using adaptive cross-attention with transformer encoder for scene graphs
Juyong Song and Sunghyun Choi · 2021
Later among the works it cites.
Boundary proposal network for two-stage natural language video localization
Shaoning Xiao, Long Chen, Songyang Zhang, Wei Ji, Jian Shao, Lu Ye, and Jun Xiao · 2021
Later among the works it cites.
Cross-category video highlight detection via set-based learning
Minghao Xu, Hang Wang, Bingbing Ni, Riheng Zhu, Zhenbang Sun, and Changhu Wang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese · 2019
Cited alongside, same era.
Language-driven temporal activity localization: A semantic matching reinforcement learning model
Weining Wang, Yan Huang, and Liang Wang · 2019
Cited alongside, same era.
Less is more: Learning highlight detection from video duration
Bo Xiong, Yannis Kalantidis, Deepti Ghadiyaram, and Kristen Grauman · 2019
Cited alongside, same era.
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan A Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko · 2019
Cited alongside, same era.
Semantic conditioned dynamic modulation for temporal sentence grounding in videos
Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu · 2019
Cited alongside, same era.
Learning modality interaction for temporal sentence localization and event captioning in videos
Shaoxiang Chen, Wenhao Jiang, Wei Liu, and Yu-Gang Jiang · 2020
Cited alongside, same era.
Tvr: A large-scale dataset for video-subtitle moment retrieval
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Cited alongside, same era.
Temporal cue guided video highlight detection with low-rank audio-visual fusion
Qinghao Ye, Xiyue Shen, Yuan Gao, Zirui Wang, Qi Bi, Ping Li, and Guang Yang · 2021
Later among the works it cites.
Conquer: Contextual query-aware ranking for video corpus moment retrieval
Hou Zhijian, Ngo Chong-Wah, and Wing-Kwong Chan · 2021
Later among the works it cites.
Contrastive learning for unsupervised video highlight detection
Taivanbat Badamdorj, Mrigank Rochan, Yang Wang, and Li Cheng · 2022
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
Learning audio-video modalities from image captions
Arsha Nagrani, Paul Hongsuck Seo, Bryan Seybold, Anja Hauth, Santiago Manen, Chen Sun, and Cordelia Schmid · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
Rethinking the video sampling and reasoning strategies for temporal sentence grounding
Jiahao Zhu, Daizong Liu, Pan Zhou, Xing Di, Yu Cheng, Song Yang, Wenzheng Xu, Zichuan Xu, Yao Wan, Lichao Sun, et al · 2022
Later among the works it cites.
Rethinking the video sampling and reasoning strategies for temporal sentence grounding
Jiahao Zhu, Daizong Liu, Pan Zhou, Xing Di, Yu Cheng, Song Yang, Wenzheng Xu, Zichuan Xu, Yao Wan, Lichao Sun, et al · 2022
Later among the works it cites.
Localizing moments in long video via multimodal guidance
Wayner Barrios, Mattia Soldan, Fabian Caba Heilbron, Alberto Mario Ceballos-Arroyo, and Bernard Ghanem · 2023
Closest in time.
Knowing where to focus: Event-aware transformer for video grounding
Jinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon, and Kwanghoon Sohn · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Univtg: Towards unified video-language temporal grounding
Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, and Mike Zheng Shou · 2023
Closest in time.
Llavilo: Boosting video moment retrieval via adapter-based multimodal modeling
Kaijing Ma, Xianghao Zang, Zerun Feng, Han Fang, Chao Ban, Yuhan Wei, Zhongjiang He, Yongxiang Li, and Hao Sun · 2023
Closest in time.
Query-dependent video representation for moment retrieval and highlight detection
WonJun Moon, Sangeek Hyun, SangUk Park, Dongchan Park, and Jae-Pil Heo · 2023
Closest in time.
Mh-detr: Video moment and highlight detection with cross-modal transformer
Yifang Xu, Yunzhuo Sun, Yang Li, Yilei Shi, Xiaoxiang Zhu, and Sidan Du · 2023
Closest in time.
Unloc: A unified framework for video localization tasks
Shen Yan, Xuehan Xiong, Arsha Nagrani, Anurag Arnab, Zhonghao Wang, Weina Ge, David Ross, and Cordelia Schmid · 2023
Closest in time.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal · 2023
Closest in time.