Fetching the paper…
Reading the bibliography…
Recently, there has been a lot of interest in automatically generating descriptions for an image.
Cognitive Psychology
U. Neisser · 1967
Earlier work this paper cites.
Selective attention and the organization of visual information
J. Duncan · 1984
Earlier work this paper cites.
Parallel versus serial processing: new vistas on the distributed organization of the visual system
J. Bulliera and L. G. Nowakb · 1995
Earlier work this paper cites.
Inferotemporal cortex and object vision
K. Tanaka · 1996
Earlier work this paper cites.
Integrated model of visual processing
J. Bullier · 2001
Earlier work this paper cites.
Dynamic predictions: Oscillations and synchrony in top-down processing
A. K. Engel, P. Fries, and W. Singer · 2001
Earlier work this paper cites.
The neural basis of perceptual learning
C. D. Gilbert, M. Sigman, and R. E. Crist · 2001
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
A cortical mechanism for triggering top-down facilitation in visual object recognition
M. Bar · 2003
Earlier work this paper cites.
Accurate unlexicalized parsing
D. Klein and C. D. Manning · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Torch7: A matlab-like environment for machine learning
R. Collobert, K. Kavukcuoglu, and C. Farabet · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Earlier work this paper cites.
Corpus-guided sentence generation of natural images
Y. Yang, C. L. Teo, H. Daumé, III, and Y. Aloimonos · 2011
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé, III · 2012
Cited alongside, same era.
Image description using visual dependency representations
D. Elliott and F. Keller · 2013
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Cited alongside, same era.
Babytalk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Later among the works it cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. N. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and F. Li · 2015
Later among the works it cites.
R. Lebret, P. H. O. Pinheiro, and R. Collobert · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Cho, B. van Merrienboer, Ç. Gülçehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Meteor universal: Language specific translation evaluation for any target language
M. Denkowski and A. Lavie · 2014
Cited alongside, same era.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
The Stanford CoreNLP natural language processing toolkit
C. D. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. J. Bethard, and D. McClosky · 2014
Cited alongside, same era.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. E. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Later among the works it cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Later among the works it cites.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Later among the works it cites.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Later among the works it cites.
phi-lstm: A phrase-based hierarchical LSTM model for image captioning
Y. H. Tan and C. S. Chan · 2016
Later among the works it cites.
Show and tell: Lessons learned from the 2015 MSCOCO image captioning challenge
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2016
Later among the works it cites.
What value do explicit high level concepts have in vision to language problems
Q. Wu, C. Shen, L. Liu, A. Dick, and A. van den Hengel · 2016
Later among the works it cites.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Later among the works it cites.