Fetching the paper…
Reading the bibliography…
Motivated by the recent progress in generative models, we introduce a model that generates images from natural language descriptions.
Information processing in dynamical systems: foundations of harmony theory
Smolensky, Paul · 1986
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Gers, Felix, Schmidhuber, Jürgen, and Cummins, Fred · 2000
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang, Zhou, Bovik, Alan C., Sheikh, Hamid R., and Simoncelli, Eero P · 2004
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, Geoffrey E., Osindero, Simon, and Teh, Yee Whye · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, Alex · 2009
Earlier work this paper cites.
Deep boltzmann machines
Salakhutdinov, Ruslan and Hinton, Geoffrey E · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, Michael and Hyvärinen, Aapo · 2010
Earlier work this paper cites.
Theano: new features and speed improvements
Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian J., Bergeron, Arnaud, Bouchard, Nicolas, Warde-Farley, David, and Bengio, Yoshua · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
Hybrid speech recognition with deep bidirectional LSTM
Graves, A., Jaitly, N., and Mohamed, A.-r · 2013
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gülçehre, Ç., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Cited alongside, same era.
Generative adversarial nets
Goodfellow, Ian J., Pouget-Abadie, Jean, Mirza, Mehdi, Xu, Bing, Warde-Farley, David, Ozair, Sherjil, Courville, Aaron C., and Bengio, Yoshua · 2014
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, Diederik P. and Welling, Max · 2014
Cited alongside, same era.
Microsoft COCO: Common objects in context
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Closest in time.
Deep generative image models using a laplacian pyramid of adversarial networks
Denton, Emily L., Chintala, Soumith, Szlam, Arthur, and Fergus, Robert · 2015
Closest in time.
DRAW: A recurrent neural network for image generation
Gregor, Karol, Danihelka, Ivo, Graves, Alex, and Wierstra, Daan · 2015
Closest in time.
Deep visual-semantic alignments for generating image descriptions
Karpathy, Andrej and Li, Fei-Fei · 2015
Closest in time.
Skip-thought vectors
Kiros, Ryan, Zhu, Yukun, Salakhutdinov, Ruslan, Zemel, Richard S., Torralba, Antonio, Urtasun, Raquel, and Fidler, Sanja · 2015
Closest in time.
Very deep convolutional networks for large-scale image recognition
Simonyan, Karen and Zisserman, Andrew · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stochastic backpropagation and variational inference in deep latent gaussian models
Rezende, Danilo J., Mohamed, Shakir, and Wierstra, Daan · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V · 2014
Cited alongside, same era.
Data generation as sequential decision making
Bachman, Philip and Precup, Doina · 2015
Cited alongside, same era.
Multimodal neural language models
Kiros, R., Salakhutdinov, R., and Zemel, R
Cited in the paper.
Unifying visual-semantic embeddings with multimodal neural language models
Kiros, Ryan, Salakhutdinov, Ruslan, and Zemel, Richard S
Cited in the paper.
Unsupervised learning of video representations using LSTMs
Srivastava, Nitish, Mansimov, Elman, and Salakhutdinov, Ruslan · 2015
Closest in time.
Show and tell: A neural image caption generator
Vinyals, Oriol, Toshev, Alexander, Bengio, Samy, and Erhan, Dumitru · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Cho, Kyunghyun, Courville, Aaron C., Salakhutdinov, Ruslan, Zemel, Richard S., and Bengio, Yoshua · 2015
Closest in time.