Fetching the paper…
Reading the bibliography…
Inspired by recent work in machine translation and object detection, we introduce an attention based model that automatically learns to describe the content of images.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
The dynamic representation of scenes
Rensink, Ronald A · 2000
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Weaver, Lex and Tao, Nigel · 2001
Earlier work this paper cites.
Control of goal-directed and stimulus-driven attention in the brain
Corbetta, Maurizio and Shulman, Gordon L · 2002
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, James, Breuleux, Olivier, Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Desjardins, Guillaume, Turian, Joseph, Warde-Farley, David, and Bengio, Yoshua · 2010
Earlier work this paper cites.
Learning to combine foveal glimpses with a third-order boltzmann machine
Larochelle, Hugo and Hinton, Geoffrey E · 2010
Earlier work this paper cites.
Composing simple image descriptions using web-scale n-grams
Li, Siming, Kulkarni, Girish, Berg, Tamara L, Berg, Alexander C, and Choi, Yejin · 2011
Earlier work this paper cites.
Corpus-guided sentence generation of natural images
Yang, Yezhou, Teo, Ching Lik, Daumé III, Hal, and Aloimonos, Yiannis · 2011
Earlier work this paper cites.
Theano: new features and speed improvements
Bastien, Frederic, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian, Bergeron, Arnaud, Bouchard, Nicolas, Warde-Farley, David, and Bengio, Yoshua · 2012
Earlier work this paper cites.
Learning where to attend with deep architectures for image tracking
Denil, Misha, Bazzani, Loris, Larochelle, Hugo, and de Freitas, Nando · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey · 2012
Earlier work this paper cites.
Collective generation of natural image descriptions
Kuznetsova, Polina, Ordonez, Vicente, Berg, Alexander C, Berg, Tamara L, and Choi, Yejin · 2012
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
Mitchell, Margaret, Han, Xufeng, Dodge, Jesse, Mensch, Alyssa, Goyal, Amit, Berg, Alex, Yamaguchi, Kota, Berg, Tamara, Stratos, Karl, and Daumé III, Hal · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Snoek, Jasper, Larochelle, Hugo, and Adams, Ryan P · 2012
Earlier work this paper cites.
Lecture 6.5 - rmsprop
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Cited alongside, same era.
Image description using visual dependency representations
Elliott, Desmond and Keller, Frank · 2013
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
Hodosh, Micah, Young, Peter, and Hockenmaier, Julia · 2013
Cited alongside, same era.
Babytalk: Understanding and generating simple image descriptions
Kulkarni, Girish, Premraj, Visruth, Ordonez, Vicente, Dhar, Sagnik, Li, Siming, Choi, Yejin, Berg, Alexander C, and Berg, Tamara L · 2013
Cited alongside, same era.
Multiple object recognition with visual attention
Ba, Jimmy Lei, Mnih, Volodymyr, and Kavukcuoglu, Koray · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Microsoft coco: Common objects in context
Lin, Tsung-Yi, Maire, Michael, Belongie, Serge, Hays, James, Perona, Pietro, Ramanan, Deva, Dollár, Piotr, and Zitnick, C Lawrence · 2014
Later among the works it cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Mao, Junhua, Xu, Wei, Yang, Yi, Wang, Jiang, and Yuille, Alan · 2014
Later among the works it cites.
Recurrent models of visual attention
Mnih, Volodymyr, Hees, Nicolas, Graves, Alex, and Kavukcuoglu, Koray · 2014
Later among the works it cites.
How to construct deep recurrent neural networks
Pascanu, Razvan, Gulcehre, Caglar, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge, 2014
Russakovsky, Olga, Deng, Jia, Su, Hao, Krause, Jonathan, Satheesh, Sanjeev, Ma, Sean, Huang, Zhiheng, Karpathy, Andrej, Khosla, Aditya, Bernstein, Michael, Berg, Alexander C., and Fei-Fei, Li · 2014
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The dropout learning algorithm
Baldi, Pierre and Sadowski, Peter · 2014
Cited alongside, same era.
Learning a recurrent visual representation for image caption generation
Chen, Xinlei and Zitnick, C Lawrence · 2014
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, Kyunghyun, van Merrienboer, Bart, Gulcehre, Caglar, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Cited alongside, same era.
Meteor universal: Language specific translation evaluation for any target language
Denkowski, Michael and Lavie, Alon · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, Jeff, Hendrikcs, Lisa Anne, Guadarrama, Segio, Rohrbach, Marcus, Venugopalan, Subhashini, Saenko, Kate, and Darrell, Trevor · 2014
Cited alongside, same era.
From captions to visual concepts and back
Fang, Hao, Gupta, Saurabh, Iandola, Forrest, Srivastava, Rupesh, Deng, Li, Dollár, Piotr, Gao, Jianfeng, He, Xiaodong, Mitchell, Margaret, Platt, John, et al · 2014
Cited alongside, same era.
Simonyan, K. and Zisserman, A · 2014
Later among the works it cites.
Input warping for bayesian optimization of non-stationary functions
Snoek, Jasper, Swersky, Kevin, Zemel, Richard S, and Adams, Ryan P · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, Nitish, Hinton, Geoffrey, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc VV · 2014
Later among the works it cites.
Going deeper with convolutions
Szegedy, Christian, Liu, Wei, Jia, Yangqing, Sermanet, Pierre, Reed, Scott, Anguelov, Dragomir, Erhan, Dumitru, Vanhoucke, Vincent, and Rabinovich, Andrew · 2014
Later among the works it cites.
Learning generative models with visual attention
Tang, Yichuan, Srivastava, Nitish, and Salakhutdinov, Ruslan R · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
Vinyals, Oriol, Toshev, Alexander, Bengio, Samy, and Erhan, Dumitru · 2014
Later among the works it cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, Peter, Lai, Alice, Hodosh, Micah, and Hockenmaier, Julia · 2014
Later among the works it cites.
Recurrent neural network regularization
Zaremba, Wojciech, Sutskever, Ilya, and Vinyals, Oriol · 2014
Later among the works it cites.