Fetching the paper…
Reading the bibliography…
Generating an image from a given text description has two goals: visual realism and semantic consistency.
Unpaired image-to-image translation using cycle-consistent adversarial networks
J. Zhu, T. Park, P. Isola, and A. A. Efros · 1903
Earlier work this paper cites.
Saccade target selection and object recognition: Evidence for a common attentional mechanism
H. Deubel and W. X. Schneider · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
Top-down control of visual attention in object detection
A. Oliva, A. Torralba, M. S. Castelhano, and J. M. Henderson · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
M.-T. Luong, H. Pham, and C. D. Manning · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Earlier work this paper cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
O. Firat, K. Cho, and Y. Bengio · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Perceptual losses for real-time style transfer and super-resolution
J. Johnson, A. Alahi, and L. Fei-Fei · 2016
Cited alongside, same era.
Generative adversarial text to image synthesis
S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee · 2016
Cited alongside, same era.
Learning what and where to draw
S. E. Reed, Z. Akata, S. Mohan, S. Tenka, B. Schiele, and H. Lee · 2016
Cited alongside, same era.
Improved techniques for training gans
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Cited alongside, same era.
Stackgan++: Realistic image synthesis with stacked generative adversarial networks
H. Zhang, T. Xu, H. Li, S. Zhang, X. Wang, X. Huang, and D. Metaxas · 2017
Later among the works it cites.
Augmented cyclegan: Learning many-to-many mappings from unpaired data
A. Almahairi, S. Rajeswar, A. Sordoni, P. Bachman, and A. C. Courville · 2018
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Later among the works it cites.
Groupcap: Group-based image captioning with structured relevance and diversity constraints
F. Chen, R. Ji, X. Sun, Y. Wu, and J. Su · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Cited alongside, same era.
Video captioning with attention-based lstm and semantic consistency
L. Gao, Z. Guo, H. Zhang, X. Xu, and H. T. Shen · 2017
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
P. Isola, J. Zhu, T. Zhou, and A. A. Efros · 2017
Cited alongside, same era.
Plug & play generative networks: Conditional iterative generation of images in latent space
A. Nguyen, J. Clune, Y. Bengio, A. Dosovitskiy, and J. Yosinski · 2017
Cited alongside, same era.
Unsupervised cross-domain image generation
Y. Taigman, A. Polyak, and L. Wolf · 2017
Cited alongside, same era.
Later among the works it cites.
Inferring semantic layout for hierarchical text-to-image synthesis
S. Hong, D. Yang, J. Choi, and H. Lee · 2018
Later among the works it cites.
Decidenet: Counting varying density crowds through attention guided detection and density estimation
J. Liu, C. Gao, D. Meng, and A. G. Hauptmann · 2018
Later among the works it cites.
Exploring human-like attention supervision in visual question answering
T. Qiao, J. Dong, and D. Xu · 2018
Later among the works it cites.
Bidirectional attentive fusion with context gating for dense video captioning
J. Wang, W. Jiang, L. Ma, W. Liu, and Y. Xu · 2018
Later among the works it cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He · 2018
Later among the works it cites.
Generative image inpainting with contextual attention
J. Yu, Z. Lin, J. Yang, X. Shen, X. Lu, and T. S. Huang · 2018
Later among the works it cites.
Progressive attention guided recurrent network for salient object detection
X. Zhang, T. Wang, J. Qi, H. Lu, and G. Wang · 2018
Later among the works it cites.
Photographic text-to-image synthesis with a hierarchically-nested adversarial network
Z. Zhang, Y. Xie, and L. Yang · 2018
Later among the works it cites.
Ancient painting to natural image: A new solution for painting processing
T. Qiao, W. Zhang, M. Zhang, Z. Ma, and D. Xu · 2019
Closest in time.