Fetching the paper…
Reading the bibliography…
Automatically generating a natural language description of an image has attracted interests recently both because of its importance in practical applications and because it connects two major artificial intelligence fields: computer vision and natural language processing.
Shifts in selective visual attention: towards the underlying neural circuitry
C. Koch and S. Ullman · 1987
Earlier work this paper cites.
A feedback model of visual attention
M. W. Spratling and M. H. Johnson · 2004
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Learning to combine foveal glimpses with a third-order boltzmann machine
H. Larochelle and G. E. Hinton · 2010
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi · 2011
Earlier work this paper cites.
Learning where to attend with deep architectures for image tracking
M. Denil, L. Bazzani, H. Larochelle, and N. de Freitas · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Collective generation of natural image descriptions
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2012
Earlier work this paper cites.
Lecture 6.5 - rmsprop, coursera: Neural networks for machine learning
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Image description using visual dependency representations
D. Elliott and F. Keller · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Deep convolutional ranking for multilabel image annotation
Y. Gong, Y. Jia, T. Leung, A. Toshev, and S. Ioffe · 2014
Cited alongside, same era.
Improving image-sentence embeddings using large weakly annotated photo collections
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik · 2014
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille · 2014
Cited alongside, same era.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, et al · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Later among the works it cites.
On the relationship between visual attributes and convolutional networks
V. Escorcia, J. C. Niebles, and B. Ghanem · 2015
Later among the works it cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al · 2015
Later among the works it cites.
Draw: A recurrent neural network for image generation
K. Gregor, I. Danihelka, A. Graves, and D. Wierstra · 2015
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Pennington, R. Socher, and C. D. Manning · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Learning generative models with visual attention
Y. Tang, N. Srivastava, and R. R. Salakhutdinov · 2014
Cited alongside, same era.
Multiple object recognition with visual attention
J. Ba, V. Mnih, and K. Kavukcuoglu · 2015
Cited alongside, same era.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Cited alongside, same era.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2015
Cited alongside, same era.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Cited alongside, same era.
R. Lebret, P. O. Pinheiro, and R. Collobert · 2015
Later among the works it cites.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Later among the works it cites.
Learning like a child: Fast novel visual concept learning from sentence descriptions of images
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Later among the works it cites.
Conceptlearner: Discovering visual concepts from weakly labeled image collections
B. Zhou, V. Jagadeesh, and R. Piramuthu · 2015
Later among the works it cites.