Fetching the paper…
Reading the bibliography…
Detecting customized moments and highlights from videos given natural language (NL) user queries is an important but under-studied topic.
The watershed transform: Definitions, algorithms and parallelization strategies
Jos BTM Roerdink and Arnold Meijster · 2000
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Query sensitive dynamic web video thumbnail generation
Chunxi Liu, Qingming Huang, and Shuqiang Jiang · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Ranking domain-specific highlights by analyzing edited videos
Min Sun, Ali Farhadi, and Steve Seitz · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Multi-task deep visual-semantic embedding for video thumbnail selection
Wu Liu, Tao Mei, Yongdong Zhang, Cherry Che, and Jiebo Luo · 2015
Earlier work this paper cites.
Video2gif: Automatic generation of animated gifs from video
Michael Gygli, Yale Song, and Liangliang Cao · 2016
Earlier work this paper cites.
To click or not to click: Automatic selection of beautiful thumbnails from videos
Yale Song, Miriam Redi, Jordi Vallmitjana, and Alejandro Jaimes · 2016
Earlier work this paper cites.
End-to-end people detection in crowded scenes
Russell Stewart, Mykhaylo Andriluka, and Andrew Y Ng · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin · 2017
Cited alongside, same era.
From lifestyle vlogs to everyday interactions
David F Fouhey, Wei-cheng Kuo, Alexei A Efros, and Jitendra Malik · 2018
Cited alongside, same era.
Phd-gifs: personalized highlight detection for automatic gif creation
Ana Garcia del Molino and Michael Gygli · 2018
Cited alongside, same era.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Cited alongside, same era.
Find and focus: Retrieve and localize video events with natural language queries
Dian Shao, Yu Xiong, Yue Zhao, Qingqiu Huang, Yu Qiao, and Dahua Lin · 2018
Cited alongside, same era.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim · 2018
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
Tvqa+: Spatio-temporal grounding for video question answering
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Later among the works it cites.
Tvr: A large-scale dataset for video-subtitle moment retrieval
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Later among the works it cites.
What is more likely to happen next? video-and-language future event prediction
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Later among the works it cites.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention augmented convolutional networks
Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V Le · 2019
Cited alongside, same era.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
Excl: Extractive clip localization using natural language descriptions
Soham Ghosh, Anuva Agarwal, Zarana Parekh, and Alexander Hauptmann · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Later among the works it cites.
Adaptive video highlight detection by learning from user history
Mrigank Rochan, Mahesh Kumar Krishna Reddy, Linwei Ye, and Yang Wang · 2020
Later among the works it cites.
A hierarchical multi-modal encoder for moment localization in video corpus
Bowen Zhang, Hexiang Hu, Joonseok Lee, Ming Zhao, Sheide Chammas, Vihan Jain, Eugene Ie, and Fei Sha · 2020
Later among the works it cites.
Span-based localizing network for natural language video localization
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou · 2020
Later among the works it cites.
Learning 2d temporal adjacent networks for moment localization with natural language
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo · 2020
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Closest in time.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Closest in time.
Interventional video grounding with dual contrastive learning
Guoshun Nan, Rui Qiao, Yao Xiao, Jun Liu, Sicong Leng, Hao Zhang, and Wei Lu · 2021
Closest in time.
Activity graph transformer for temporal action localization
Megha Nawhal and Greg Mori · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
Temporal query networks for fine-grained video understanding
Chuhan Zhang, Ankush Gupta, and Andrew Zisserman · 2021
Closest in time.