Fetching the paper…
Reading the bibliography…
Automatic synthesis of realistic images from text would be interesting and useful, but current AI systems are still far from this goal.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Describing objects by their attributes
Farhadi, A., Endres, I., Hoiem, D., and Forsyth, D · 2009
Earlier work this paper cites.
Attribute and simile classifiers for face verification
Kumar, N., Berg, A. C., Belhumeur, P. N., and Nayar, S. K · 2009
Earlier work this paper cites.
Multimodal deep learning
Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A. Y · 2011
Earlier work this paper cites.
Relative attributes
Parikh, D. and Grauman, K · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S · 2011
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
Srivastava, N. and Salakhutdinov, R. R · 2012
Earlier work this paper cites.
Better mixing via deep representations
Bengio, Y., Mesnil, G., Dauphin, Y., and Rifai, S · 2013
Earlier work this paper cites.
Transductive multi-view embedding for zero-shot recognition and annotation
Fu, Y., Hospedales, T. M., Xiang, T., Fu, Z., and Gong, S · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Kiros, R., Salakhutdinov, R., and Zemel, R. S · 2014
Earlier work this paper cites.
Attribute-based classification for zero-shot visual object categorization
Lampert, C. H., Nickisch, H., and Harmeling, S · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mirza, M. and Osindero, S · 2014
Cited alongside, same era.
Learning to disentangle factors of variation with manifold interaction
Reed, S., Sohn, K., Zhang, Y., and Lee, H · 2014
Cited alongside, same era.
Improved multimodal deep learning with variation of information
Sohn, K., Shang, W., and Lee, H · 2014
Cited alongside, same era.
Evaluation of Output Embeddings for Fine-Grained Image Classification
Akata, Z., Reed, S., Walter, D., Lee, H., and Schiele, B · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Ba, J. and Kingma, D · 2015
Cited alongside, same era.
Deep generative image models using a laplacian pyramid of adversarial networks
Denton, E. L., Chintala, S., Fergus, R., et al · 2015
Deep visual analogy-making
Reed, S., Zhang, Y., Zhang, Y., and Lee, H · 2015
Later among the works it cites.
Exploring models and data for image question answering
Ren, M., Kiros, R., and Zemel, R · 2015
Later among the works it cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D · 2015
Later among the works it cites.
Explicit knowledge-based reasoning for visual question answering
Wang, P., Wu, Q., Shen, C., Hengel, A. v. d., and Dick, A · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Hendricks, L. A., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., and Darrell, T · 2015
Cited alongside, same era.
Learning to generate chairs with convolutional neural networks
Dosovitskiy, A., Tobias Springenberg, J., and Brox, T · 2015
Cited alongside, same era.
Conditional generative adversarial nets for convolutional face generation
Gauthier, J · 2015
Cited alongside, same era.
Draw: A recurrent neural network for image generation
Gregor, K., Danihelka, I., Graves, A., Rezende, D., and Wierstra, D · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Li, F · 2015
Cited alongside, same era.
Later among the works it cites.
Attribute2image: Conditional image generation from visual attributes
Yan, X., Yang, J., Sohn, K., and Lee, H · 2015
Later among the works it cites.
Weakly-supervised disentangling with recurrent transformations for 3d view synthesis
Yang, J., Reed, S., Yang, M.-H., and Lee, H · 2015
Later among the works it cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Later among the works it cites.
Generating images from captions with attention
Mansimov, E., Parisotto, E., Ba, J. L., and Salakhutdinov, R · 2016
Closest in time.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A., Metz, L., and Chintala, S · 2016
Closest in time.
Learning deep representations for fine-grained visual descriptions
Reed, S., Akata, Z., Lee, H., and Schiele, B · 2016
Closest in time.