Fetching the paper…
Reading the bibliography…
Text-to-image generation models that generate images based on prompt descriptions have attracted an increasing amount of attention during the past few months.
A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise
Martin Ester, Hans-Peter Kriegel, Jörg Sander, and Xiaowei Xu · 1996
Earlier work this paper cites.
NLTK: The Natural Language Toolkit
Steven Bird and Edward Loper · 2004
Earlier work this paper cites.
Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
Hodosh Micah, Young Peter, and Hockenmaier Julia · 2013
Earlier work this paper cites.
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Diffusion-Convolutional Neural Networks
James Atwood and Don Towsley · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Generative Adversarial Text to Image Synthesis
Scott E. Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Photographic Image Synthesis with Cascaded Refinement Networks
Qifeng Chen and Vladlen Koltun · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks
Han Zhang, Tao Xu, and Hongsheng Li · 2017
Earlier work this paper cites.
Semi-supervised FusedGAN for Conditional Image Generation
Navaneeth Bodla, Gang Hua, and Rama Chellappa · 2018
Earlier work this paper cites.
Learning to See in the Dark
Chen Chen, Qifeng Chen, Jia Xu, and Vladlen Koltun · 2018
Cited alongside, same era.
StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation
Yunjey Choi, Min-Je Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo · 2018
Cited alongside, same era.
Second-Order Attention Network for Single Image Super-Resolution
Tao Dai, Jianrui Cai, Yongbing Zhang, Shu-Tao Xia, and Lei Zhang · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Dual Adversarial Inference for Text-to-Image Synthesis
Qicheng Lao, Mohammad Havaei, Ahmad Pesaranghader, Francis Dutil, Lisa Di-Jorio, and Thomas Fevens · 2019
Cited alongside, same era.
Diverse Image Synthesis From Semantic Layouts via Conditional IMLE
Concadia: Tackling Image Accessibility with Descriptive Texts and Context
Kreiss Elisa, Goodman Noah D, and Potts Christopher · 2021
Later among the works it cites.
Towards Discovery and Attribution of Open-World GAN Generated Images
Sharath Girish, Saksham Suri, Sai Saketh Rambhatla, and Abhinav Shrivastava · 2021
Later among the works it cites.
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ke Li, Tianhao Zhang, and Jitendra Malik · 2019
Cited alongside, same era.
Object-Driven Text-To-Image Synthesis via Adversarial Training
Wenbo Li, Pengchuan Zhang, Lei Zhang, Qiuyuan Huang, Xiaodong He, Siwei Lyu, and Jianfeng Gao · 2019
Cited alongside, same era.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Attributing Fake Images to GANs: Learning and Analyzing GAN Fingerprints
Ning Yu, Larry Davis, and Mario Fritz · 2019
Cited alongside, same era.
Detecting and Simulating Artifacts in GAN Fake Images
Xu Zhang, Svebor Karaman, and Shih-Fu Chang · 2019
Cited alongside, same era.
Efficient Neural Architecture for Text-to-Image Synthesis
Douglas M. Souza, Jonatas Wehrmann, and Duncan D. Ruiz · 2020
Cited alongside, same era.
CNN-Generated Images Are Surprisingly Easy to Spot… for Now
Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A. Efros · 2020
Cited alongside, same era.
Later among the works it cites.
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Later among the works it cites.
Cross-Modal Contrastive Learning for Text-to-Image Generation
Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang · 2021
Later among the works it cites.
Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge
Pierre L. Dognin, Igor Melnyk, Youssef Mroueh, Inkit Padhi, Mattia Rigotti, Jarret Ross, Yair Schiff, Richard A. Young, and Brian Belgodere · 2022
Closest in time.
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi · 2022
Closest in time.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Closest in time.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi · 2022
Closest in time.
Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning
Shibani Santurkar, Yann Dubois, Rohan Taori, Percy Liang, and Tatsunori Hashimoto · 2022
Closest in time.