Fetching the paper…
Reading the bibliography…
Joint modeling of language and vision has been drawing increasing interest.
Relations between two sets of variates
Harold Hotteling · 1936
Earlier work this paper cites.
Numerical methods for computing angles between linear subspaces
Gene H. Golub, He Bjlirck, and Gene H. Golub · 1973
Earlier work this paper cites.
Canonical ridge and econometrics of joint production
H.D. Vinod · 1976
Earlier work this paper cites.
The truncated svd as a method for regularization
Per C. Hansen · 1986
Earlier work this paper cites.
Springer, 1995
Gene H. Golub and Hongyuan Zha · 1995
Earlier work this paper cites.
Language Modeling for Information Retrieval
John Lafferty and ChengXiang Zhai · 2003
Earlier work this paper cites.
Bhattacharyya and expected likelihood kernels
Tony Jebara and Risi Kondor · 2003
Earlier work this paper cites.
A hilbert space embedding for distributions
Alex Smola, Arthur Gretton, Le Song, and Bernhard Schölkopf · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Universal kernels on non-standard input spaces
Andreas Christmann and Ingo Steinwart · 2010
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
Richard Socher, Milind Ganjoo, Hamsa Sridhar, Osbert Bastani, Christopher D. Manning, and Andrew Y. Ng · 2013
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Cited alongside, same era.
Explain images with multimodal recurrent neural networks
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L. Yuille · 2014
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2014
Cited alongside, same era.
Improving image-sentence embeddings using large weakly annotated photo collections
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik · 2014
From captions to visual concepts and back
H. Fang, S. Gupta, F. N. Iandola, R.K Srivastava, L. Deng, P. Dollár, J. Gao, X. He, Margaret. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Closest in time.
Associating neural word embeddings with deep image representations using fisher vectors
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf · 2015
Closest in time.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2015
Closest in time.
Cross-domain matching for bag-of-words data via kernel embeddings of latent distributions
Yuya Yoshikawa, Tomoharu Iwata, Hiroshi Sawada, and Takeshi Yamada · 2015
Closest in time.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Closest in time.
Order-embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li · 2015
Cited alongside, same era.
Closest in time.
Word representations via gaussian embedding
Luke Vilnis and Andrew McCallum · 2015
Closest in time.
Learning deep structure-preserving image-text embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Closest in time.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Closest in time.