Fetching the paper…
Reading the bibliography…
Identifying a short segment in a long video that semantically matches a text query is a challenging task that has important application potentials in language-based video search, browsing, and navigation.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko · 2014
Earlier work this paper cites.
Sequence to sequence – video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
Ran Xu, Caiming Xiong, Wei Chen, and Jason J. Corso · 2015
Earlier work this paper cites.
Word2VisualVec: Cross-media retrieval by visual feature prediction
Jianfeng Dong, Xirong Li, and Cees G. M. Snoek · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Cisco global cloud index: Forecast and methodology, 2016–2021
Cisco Visual Networking · 2016
Earlier work this paper cites.
Jointly modeling embedding and translation to bridge video and language
Yingwei Pan, Tao Mei, Ting Yao, Houqiang Li, and Yong Rui · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Vse++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2017
Cited alongside, same era.
TALL: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Cited alongside, same era.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell · 2019
Later among the works it cites.
DEBUG: A dense bottom-up grounding approach for natural language video localization
Chujie Lu, Long Chen, Chilie Tan, Xiaolin Li, and Jun Xiao · 2019
Later among the works it cites.
VilBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
VideoBERT: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, undefinedukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Localizing moments in video with temporal language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2018
Cited alongside, same era.
Cross-modal moment localization in videos
Meng Liu, Xiang Wang, Liqiang Nie, Qi Tian, Baoquan Chen, and Tat-Seng Chua · 2018
Cited alongside, same era.
Learning joint embedding with multimodal cues for cross-modal video-text retrieval
Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, and Amit K. Roy-Chowdhury · 2018
Cited alongside, same era.
Find and focus: Retrieve and localize video events with natural language queries
Dian Shao, Yu Xiong, Yue Zhao, Qingqiu Huang, Yu Qiao, and Dahua Lin · 2018
Cited alongside, same era.
Semantic conditioned dynamic modulation for temporal sentence grounding in videos
Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu · 2019
Later among the works it cites.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yitian Yuan, Tao Mei, and Wenwu Zhu · 2019
Later among the works it cites.
Learning the best pooling strategy for visual semantic embedding
Jiacheng Chen, Hexiang Hu, Hao Wu, Yuning Jiang, and Changhu Wang · 2020
Closest in time.
Rethinking the bottom-up framework for query-based video localization
Long Chen, Chujie Lu, Siliang Tang, Jun Xiao, Dong Zhang, Chilie Tan, and Xiaolin Li · 2020
Closest in time.
Entities as experts: Sparse memory access with entity supervision
Thibault Févry, Livio Baldini Soares, Nicholas FitzGerald, Eunsol Choi, and Tom Kwiatkowski · 2020
Closest in time.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy · 2020
Closest in time.
TVR: A large-scale dataset for video-subtitle moment retrieval
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Closest in time.
Hero: Hierarchical encoder for video+language omni-representation pre-training, 2020
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Closest in time.