Fetching the paper…
Reading the bibliography…
The rapid growth of video on the internet has made searching for video content using natural language queries a significant challenge.
Hierarchical mixtures of experts and the em algorithm
Michael I Jordan and Robert A Jacobs · 1994
Earlier work this paper cites.
The OpenCV Library
G. Bradski · 2000
Earlier work this paper cites.
Multi-modal information retrieval from broadcast video using ocr and speech recognition
Alexander G Hauptmann, Rong Jin, and Tobun Dorbin Ng · 2002
Earlier work this paper cites.
Topic segmentation and retrieval system for lecture videos based on spontaneous speech recognition
Natsuo Yamamoto, Jun Ogata, and Yasuo Ariki · 2003
Earlier work this paper cites.
Large-scale content-based audio retrieval from text queries
Gal Chechik, Eugene Ie, Martin Rehn, Samy Bengio, and Dick Lyon · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
The unreasonable effectiveness of data
Alon Halevy, Peter Norvig, and Fernando Pereira · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David L Chen and William B Dolan · 2011
Earlier work this paper cites.
Combining attributes and fisher vectors for efficient image retrieval
Matthijs Douze, Arnau Ramisa, and Cordelia Schmid · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Tomas Mikolov, et al · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Synthetic data and artificial neural networks for natural scene text recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V Le, Christopher D Manning, and Andrew Y Ng · 2014
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf · 2015
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
Ran Xu, Caiming Xiong, Wei Chen, and Jason J Corso · 2015
Earlier work this paper cites.
Netvlad: Cnn architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic · 2016
Earlier work this paper cites.
Word2visualvec: Image and video to sentence matching by visual feature prediction
Jianfeng Dong, Xirong Li, and Cees GM Snoek · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Ssd: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg · 2016
Cited alongside, same era.
Learning joint representations of videos and sentences with web image search
Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, and Naokazu Yokoya · 2016
Cited alongside, same era.
Jointly modeling embedding and translation to bridge video and language
Yingwei Pan, Tao Mei, Ting Yao, Houqiang Li, and Yong Rui · 2016
Cited alongside, same era.
Video retrieval using speech and text in video
N Radha · 2016
Cited alongside, same era.
Learning language-visual embedding for movie understanding with natural-language
Pixellink: Detecting scene text via instance segmentation
Dan Deng, Haifeng Liu, Xuelong Li, and Deng Cai · 2018
Later among the works it cites.
Predicting visual features from text for image and video caption retrieval
Jianfeng Dong, Xirong Li, and Cees GM Snoek · 2018
Later among the works it cites.
Massively parallel hyperparameter tuning
Liam Li, Kevin Jamieson, Afshin Rostamizadeh, Ekaterina Gonina, Moritz Hardt, Benjamin Recht, and Ameet Talwalkar · 2018
Later among the works it cites.
Synthetically supervised feature learning for scene text recognition
Yang Liu, Zhaowen Wang, Hailin Jin, and Ian Wassell · 2018
Later among the works it cites.
Exploring the limits of weakly supervised pretraining
Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Atousa Torabi, Niket Tandon, and Leonid Sigal · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Video captioning and retrieval models with semantic attention
Youngjae Yu, Hyungjin Ko, Jongwook Choi, and Gunhee Kim · 2016
Cited alongside, same era.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Cited alongside, same era.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Vse++: Improved visual-semantic embeddings
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2017
Cited alongside, same era.
Antoine Miech, Ivan Laptev, and Josef Sivic · 2018
Later among the works it cites.
Learning joint embedding with multimodal cues for cross-modal video-text retrieval
Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, and Amit K Roy-Chowdhury · 2018
Later among the works it cites.
Learnable pins: Cross-modal embeddings for person identity
Arsha Nagrani, Samuel Albanie, and Andrew Zisserman · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Later among the works it cites.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim · 2018
Later among the works it cites.
Cross-modal and hierarchical modeling of video and text
Bowen Zhang, Hexiang Hu, and Fei Sha · 2018
Later among the works it cites.
Ghostvlad for set-based face recognition
Yujie Zhong, Relja Arandjelović, and Andrew Zisserman · 2018
Later among the works it cites.
ig65m-pytorch
J. H. Daniel · 2019
Closest in time.
Dual dense encoding for zero-example video retrieval
Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, and Xun Wang · 2019
Closest in time.
Large-scale weakly-supervised pre-training for video action recognition
Deepti Ghadiyaram, Du Tran, and Dhruv Mahajan · 2019
Closest in time.
Squeeze-and-excitation networks
Jie Hu, Li Shen, Samuel Albanie, Gang Sun, and Enhua Wu · 2019
Closest in time.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Closest in time.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2019
Closest in time.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Closest in time.
Project title
Less Wright · 2019
Closest in time.
Lookahead optimizer: k steps forward, 1 step back
Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton · 2019
Closest in time.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic · 2019
Closest in time.