Fetching the paper…
Reading the bibliography…
We present a new technique for learning visual-semantic embeddings for cross-modal retrieval.
Histograms of oriented gradients for human detection
Navneet Dalal and Bill Triggs · 2005
Earlier work this paper cites.
Large margin methods for structured and interdependent output variables
Ioannis Tsochantaridis, Thorsten Joachims, Thomas Hofmann, and Yasemin Altun · 2005
Earlier work this paper cites.
Large margin optimization of ranking measures
Olivier Chapelle, Quoc Le, and Alex Smola · 2007
Earlier work this paper cites.
Learning globally-consistent local distance functions for shape-based image retrieval and classification
Andrea Frome, Yoram Singer, Fei Sha, and Jitendra Malik · 2007
Earlier work this paper cites.
Direct optimization of ranking measures
Quoc Le and Alexander Smola · 2007
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Learning structural svms with latent variables
Chun-Nam John Yu and Thorsten Joachims · 2009
Earlier work this paper cites.
Large scale online learning of image similarity through ranking
Gal Chechik, Varun Sharma, Uri Shalit, and Samy Bengio · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
Pedro F Felzenszwalb, Ross B Girshick, David McAllester, and Deva Ramanan · 2010
Earlier work this paper cites.
Ensemble of exemplar-svms for object detection and beyond
Tomasz Malisiewicz, Abhinav Gupta, and Alexei A Efros · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Tomas Mikolov, et al · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Earlier work this paper cites.
Bilingual word embeddings for phrase-based machine translation
Will Y Zou, Richard Socher, Daniel Cer, and Christopher D Manning · 2013
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Cited alongside, same era.
Learning to rank for information retrieval and natural language processing
Hang Li · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Cited alongside, same era.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V Le, Christopher D Manning, and Andrew Y Ng · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Later among the works it cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Later among the works it cites.
Order-embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2016
Later among the works it cites.
Measuring machine intelligence through visual question answering
C Lawrence Zitnick, Aishwarya Agrawal, Stanislaw Antol, Margaret Mitchell, Dhruv Batra, and Devi Parikh · 2016
Later among the works it cites.
Vqa: Visual question answering
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C Lawrence Zitnick, Devi Parikh, and Dhruv Batra · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Associating neural word embeddings with deep image representations using fisher vectors
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz · 2015
Cited alongside, same era.
Facenet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Cited alongside, same era.
Maximum-margin structured learning with deep networks for 3d human pose estimation
Sijin Li, Weichen Zhang, and Antoni B Chan
Cited in the paper.
Closest in time.
Linking image and text with 2-way nets
Aviv Eisenschtat and Lior Wolf · 2017
Closest in time.
Instance-aware image and sentence matching with selective multimodal lstm
Yan Huang, Wei Wang, and Liang Wang · 2017
Closest in time.
Dual attention networks for multimodal reasoning and matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim · 2017
Closest in time.
Sampling matters in deep embedding learning
Chao-Yuan Wu, R Manmatha, Alexander J Smola, and Philipp Krähenbühl · 2017
Closest in time.
Learning two-branch neural networks for image-text matching tasks
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2018
Closest in time.