Fetching the paper…
Reading the bibliography…
In the text-to-image generation field, recent remarkable progress in Stable Diffusion makes it possible to generate rich kinds of novel photorealistic images.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Monte Carlo sampling methods
Alexander Shapiro. 2003 · 2003
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Generative Adversarial Nets. In NeurIPS . 2672–2680
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes. In ICLR
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In ECCV , Vol. 8693. 740–755
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation. In MICCAI . 234–241
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
CVAE-GAN: Fine-Grained Image Generation through Asymmetric Training. In ICCV . 2764–2773
Jianmin Bao, Dong Chen, Fang Wen, Houqiang Li, and Gang Hua. 2017 · 2017
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In NeurIPS . 6626–6637
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks. In ICCV . 5908–5916
Han Zhang, Tao Xu, and Hongsheng Li. 2017 · 2017
Earlier work this paper cites.
Inferring Semantic Layout for Hierarchical Text-to-Image Synthesis. In CVPR . 7986–7994
Seunghoon Hong, Dingdong Yang, Jongwook Choi, and Honglak Lee. 2018 · 2018
Earlier work this paper cites.
Image Generation From Scene Graphs. In CVPR . 1219–1228
Justin Johnson, Agrim Gupta, and Li Fei-Fei. 2018 · 2018
Earlier work this paper cites.
AttnGAN: Fine-Grained Text to Image Generation With Attentional Generative Adversarial Networks. In CVPR . 1316–1324
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018 · 2018
Earlier work this paper cites.
Generating Multiple Objects at Spatially Distinct Locations. In ICLR
Tobias Hinz, Stefan Heinrich, and Stefan Wermter. 2019 · 2019
Earlier work this paper cites.
On the Spectral Bias of Neural Networks. In ICML . 5301–5310
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A. Hamprecht, Yoshua Bengio, and Aaron C. Courville. 2019 · 2019
Earlier work this paper cites.
Image Synthesis From Reconfigurable Layout and Style. In ICCV . 10530–10539
Wei Sun and Tianfu Wu. 2019 · 2019
Earlier work this paper cites.
Denoising Diffusion Probabilistic Models. In NeurIPS
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Attribute-Guided Image Generation from Layout. In BMVC
Ke Ma, Bo Zhao, and Leonid Sigal. 2020 · 2020
Earlier work this paper cites.
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. In ECCV . 405–421
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. 2020 · 2020
Earlier work this paper cites.
READ: Recursive Autoencoders for Document Layout Generation. In CVPR . 2316–2325
Akshay Gadi Patil, Omri Ben-Eliezer, Or Perel, and Hadar Averbuch-Elor. 2020 · 2020
Earlier work this paper cites.
Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains. In NeurIPS
Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. 2020 · 2020
Earlier work this paper cites.
Visual-Relation Conscious Image Generation from Structured-Text. In ECCV . 290–306
Duc Minh Vo and Akihiro Sugimoto. 2020 · 2020
Cited alongside, same era.
Diffusion Models Beat GANs on Image Synthesis. In NeurIPS . 8780–8794
Prafulla Dhariwal and Alexander Quinn Nichol. 2021 · 2021
Cited alongside, same era.
CogView: Mastering Text-to-Image Generation via Transformers. In NeurIPS . 19822–19835
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. 2021 · 2021
Cited alongside, same era.
LayoutTransformer: Layout Generation and Completion with Self-attention. In ICCV . 984–994
Kamal Gupta, Justin Lazarow, Alessandro Achille, Larry Davis, Vijay Mahadevan, and Abhinav Shrivastava. 2021 · 2021
Cited alongside, same era.
Context-Aware Layout to Image Generation With Enhanced Object Appearance. In CVPR . 15049–15058
Sen He, Wentong Liao, Michael Ying Yang, Yongxin Yang, Yi-Zhe Song, Bodo Rosenhahn, and Tao Xiang. 2021 · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022a · 2022
Later among the works it cites.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022b · 2022
Later among the works it cites.
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Constrained Graphic Layout Generation via Latent Optimization. In ACM MM . 88–96
Kotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, and Kota Yamaguchi. 2021 · 2021
Cited alongside, same era.
Dynamic modality interaction modeling for image-text retrieval. In SIGIR . 1104–1113
Leigang Qu, Meng Liu, Jianlong Wu, Zan Gao, and Liqiang Nie. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision. In ICML . 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
Zero-Shot Text-to-Image Generation. In ICML . 8821–8831
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
Calibrate Before Use: Improving Few-shot Performance of Language Models. In ICML (Proceedings of Machine Learning Research, Vol. 139) . 12697–12706
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
SpaText: Spatio-Textual Representation for Controllable Image Generation
Omri Avrahami, Thomas Hayes, Oran Gafni, Sonal Gupta, Yaniv Taigman, Devi Parikh, Dani Lischinski, Ohad Fried, and Xi Yin. 2022 · 2022
Cited alongside, same era.
MaskGIT: Masked Generative Image Transformer. In CVPR . 11305–11315
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. 2022 · 2022
Cited alongside, same era.
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. 2022 · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Later among the works it cites.
Sketch-Guided Text-to-Image Diffusion Models
Andrey Voynov, Kfir Aberman, and Daniel Cohen-Or. 2022 · 2022
Later among the works it cites.
Chain of Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. 2022 · 2022
Later among the works it cites.
Active Example Selection for In-Context Learning. In NeurIPS . 9134–9148
Yiming Zhang, Shi Feng, and Chenhao Tan. 2022 · 2022
Later among the works it cites.
MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation
Omer Bar-Tal, Lior Yariv, Yaron Lipman, and Tali Dekel. 2023 · 2023
Closest in time.
LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation
Jiaxin Cheng, Xiao Liang, Xingjian Shi, Tong He, Tianjun Xiao, and Mu Li. 2023 · 2023
Closest in time.
LayoutDM: Discrete Diffusion Model for Controllable Layout Generation
Naoto Inoue, Kotaro Kikuchi, Edgar Simo-Serra, Mayu Otani, and Kota Yamaguchi. 2023 · 2023
Closest in time.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023a · 2023
Closest in time.
GLIGEN: Open-Set Grounded Text-to-Image Generation
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. 2023b · 2023
Closest in time.
Chong Mou, Xintao Wang, Liangbin Xie, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Adding Conditional Control to Text-to-Image Diffusion Models
Lvmin Zhang and Maneesh Agrawala. 2023a · 2023
Closest in time.
Adding Conditional Control to Text-to-Image Diffusion Models
Lvmin Zhang and Maneesh Agrawala. 2023b · 2023
Closest in time.
LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
Guangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi, Ying Shan, and Xi Li. 2023 · 2023
Closest in time.