Fetching the paper…
Reading the bibliography…
We introduce a model for bidirectional retrieval of images and sentences through a multi-modal embedding of visual and natural language data.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., Haffner, P.: · 1998
Earlier work this paper cites.
Multiple instance learning with generalized support vector machines
Andrews, S., Hofmann, T., Tsochantaridis, I.: · 2002
Earlier work this paper cites.
Matching words and pictures
Barnard, K., Duygulu, P., Forsyth, D., De Freitas, N., Blei, D.M., Jordan, M.I.: · 2003
Earlier work this paper cites.
Generating typed dependency parses from phrase structure parses
De Marneffe, M.C., MacCartney, B., Manning, C.D., et al.: · 2006
Earlier work this paper cites.
Neural probabilistic language models
Bengio, Y., Schwenk, H., Senécal, J.S., Morin, F., Gauvain, J.L.: · 2006
Earlier work this paper cites.
Miles: Multiple-instance learning via embedded instance selection
Chen, Y., Bi, J., Wang, J.Z.: · 2006
Earlier work this paper cites.
Three new graphical models for statistical language modelling
Mnih, A., Hinton, G.: · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R., Weston, J.: · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: · 2009
Earlier work this paper cites.
Collecting image annotations using amazon’s mechanical turk
Rashtchian, C., Young, P., Hodosh, M., Hockenmaier, J.: · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Farhadi, A., Hejrati, M., Sadeghi, M.A., Young, P., Rashtchian, C., Hockenmaier, J., Forsyth, D.: · 2010
Earlier work this paper cites.
I2t: Image parsing to text description
Yao, B.Z., Yang, X., Lin, L., Lee, M.W., Zhu, S.C.: · 2010
Earlier work this paper cites.
Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora
Socher, R., Fei-Fei, L.: · 2010
Earlier work this paper cites.
Word representations: a simple and general method for semi-supervised learning
Turian, J., Ratinov, L., Bengio, Y.: · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., Berg, T.L.: · 2011
Cited alongside, same era.
Baby talk: Understanding and generating simple image descriptions
Kulkarni, G., Premraj, V., Dhar, S., Li, S., Choi, Y., Berg, A.C., Berg, T.L.: · 2011
Cited alongside, same era.
Corpus-guided sentence generation of natural images
Yang, Y., Teo, C.L., Daumé III, H., Aloimonos, Y.: · 2011
Cited alongside, same era.
Composing simple image descriptions using web-scale n-grams
Li, S., Kulkarni, G., Berg, T.L., Berg, A.C., Choi, Y.: · 2011
Cited alongside, same era.
Learning cross-modality similarity for multinomial data
Jia, Y., Salzmann, M., Darrell, T.: · 2011
Cited alongside, same era.
Multimodal deep learning
Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., Ng, A.Y.: · 2011
Cited alongside, same era.
Learning the visual interpretation of sentences
Zitnick, C.L., Parikh, D., Vanderwende, L.: · 2013
Later among the works it cites.
Devise: A deep visual-semantic embedding model
Frome, A., Corrado, G.S., Shlens, J., Bengio, S., Dean, J., Mikolov, T., et al.: · 2013
Later among the works it cites.
Building high-level features using large scale unsupervised learning
Le, Q.V.: · 2013
Later among the works it cites.
Visualizing and understanding convolutional neural networks
Zeiler, M.D., Fergus, R.: · 2013
Later among the works it cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: · 2013
Later among the works it cites.
Large scale visual recognition challenge 2013
Russakovsky, O., Deng, J., Krause, J., Berg, A., Fei-Fei, L.: · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parsing natural scenes and natural language with recursive neural networks
Socher, R., Lin, C.C., Manning, C., Ng, A.Y.: · 2011
Cited alongside, same era.
Midge: Generating image descriptions from computer vision detections
Mitchell, M., Han, X., Dodge, J., Mensch, A., Goyal, A., Berg, A., Yamaguchi, K., Berg, T., Stratos, K., Daumé, III, H.: · 2012
Cited alongside, same era.
Collective generation of natural image descriptions
Kuznetsova, P., Ordonez, V., Berg, A.C., Berg, T.L., Choi, Y.: · 2012
Cited alongside, same era.
Multimodal learning with deep boltzmann machines
Srivastava, N., Salakhutdinov, R.: · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G.E.: · 2012
Cited alongside, same era.
Improving word representations via global context and multiple word prototypes
Huang, E.H., Socher, R., Manning, C.D., Ng, A.Y.: · 2012
Cited alongside, same era.
Later among the works it cites.
Caffe: An open source convolutional architecture for fast feature embedding
Jia, Y.: · 2013
Later among the works it cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., Hockenmaier, J.: · 2014
Closest in time.
Multimodal neural language models
Kiros, R., Zemel, R.S., Salakhutdinov, R.: · 2014
Closest in time.
Grounded compositional semantics for finding and describing images with sentences
Socher, R., Karpathy, A., Le, Q.V., Manning, C.D., Ng, A.Y.: · 2014
Closest in time.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., Malik, J.: · 2014
Closest in time.
Overfeat: Integrated recognition, localization and detection using convolutional networks
Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., LeCun, Y.: · 2014
Closest in time.
Distributed representations of sentences and documents
Le, Q.V., Mikolov, T.: · 2014
Closest in time.