Fetching the paper…
Reading the bibliography…
The tremendous progress in neural image generation, coupled with the emergence of seemingly omnipotent vision-language models has finally enabled text-based interfaces for creating and editing images.
StackGAN++: Realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. 2018b · 1962
Earlier work this paper cites.
Poisson image editing
Patrick Pérez, Michel Gangnet, and Andrew Blake. 2003 · 2003
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database. In Proc. CVPR . Ieee, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
Generative Adversarial Nets. In Advances in Neural Information Processing Systems . 2672–2680
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In Proc. ICML . 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Generating Images from Captions with Attention
Elman Mansimov, Emilio Parisotto, Jimmy Ba, and Ruslan Salakhutdinov. 2016 · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis. In Proc. ICLR . 1060–1069
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016 · 2016
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning. In NIPS
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks. In Proc. ICCV . 5907–5915
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. 2017 · 2017
Earlier work this paper cites.
Large scale GAN training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan. 2018 · 2018
Earlier work this paper cites.
AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1316–1324
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. 2018 · 2018
Earlier work this paper cites.
Image2stylegan: How to embed images into the StyleGAN latent space?. In Proc. ICCV . 4432–4441
Rameen Abdal, Yipeng Qin, and Peter Wonka. 2019 · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4401–4410
Tero Karras, Samuli Laine, and Timo Aila. 2019 · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with VQ-VAE-2
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals. 2019 · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning . PMLR, 6105–6114
Mingxing Tan and Quoc Le. 2019 · 2019
Earlier work this paper cites.
Image2stylegan++: How to edit the embedded images?. In Proc. CVPR . 8296–8305
Rameen Abdal, Yipeng Qin, and Peter Wonka. 2020 · 2020
Earlier work this paper cites.
Semantic photo manipulation with a generative image prior
David Bau, Hendrik Strobelt, William Peebles, Jonas Wulff, Bolei Zhou, Jun-Yan Zhu, and Antonio Torralba. 2020 · 2020
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Editing Self-Image
Ohad Fried, Jennifer Jacobs, Adam Finkelstein, and Maneesh Agrawala. 2020 · 2020
Cited alongside, same era.
Denoising Diffusion Probabilistic Models. In Proc. NeurIPS
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Cited alongside, same era.
Analyzing and improving the image quality of StyleGAN. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8110–8119
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020 · 2020
Cited alongside, same era.
Denoising Diffusion Implicit Models. In International Conference on Learning Representations
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020 · 2020
Cited alongside, same era.
In-domain GAN inversion for real image editing. In Proc. ECCV . Springer, 592–608
Jiapeng Zhu, Yujun Shen, Deli Zhao, and Bolei Zhou. 2020 · 2020
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Later among the works it cites.
Score-based generative modeling in latent space
Arash Vahdat, Karsten Kreis, and Jan Kautz. 2021 · 2021
Later among the works it cites.
Weihao Xia, Yulun Zhang, Yujiu Yang, Jing-Hao Xue, Bolei Zhou, and Ming-Hsuan Yang. 2021 · 2021
Later among the works it cites.
LAFITE: Towards Language-Free Training for Text-to-Image Generation
Yufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, and Tong Sun. 2021 · 2021
Later among the works it cites.
Amazon Mechanical Turk
Amazon. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Bau, Alex Andonian, Audrey Cui, YeonHwan Park, Ali Jahanian, Aude Oliva, and Antonio Torralba. 2021 · 2021
Cited alongside, same era.
Sam Bond-Taylor, Peter Hessey, Hiroshi Sasaki, Toby P Breckon, and Chris G Willcocks. 2021 · 2021
Cited alongside, same era.
ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, 14347–14356
Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. 2021 · 2021
Cited alongside, same era.
CLIP Guided Diffusion HQ 256x256
Katherine Crowson. 2021 · 2021
Cited alongside, same era.
Diffusion models beat GANs on image synthesis
Prafulla Dhariwal and Alexander Nichol. 2021 · 2021
Cited alongside, same era.
Cogview: Mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al · 2021
Cited alongside, same era.
ImageBART: Bidirectional context with multinomial diffusion for autoregressive image synthesis
Patrick Esser, Robin Rombach, Andreas Blattmann, and Bjorn Ommer. 2021b · 2021
Cited alongside, same era.
Closest in time.
KNN-Diffusion: Image Generation via Large-Scale Retrieval
Oron Ashual, Shelly Sheynin, Adam Polyak, Uriel Singer, Oran Gafni, Eliya Nachmani, and Yaniv Taigman. 2022 · 2022
Closest in time.
SpaText: Spatio-Textual Representation for Controllable Image Generation
Omri Avrahami, Thomas Hayes, Oran Gafni, Sonal Gupta, Yaniv Taigman, Devi Parikh, Dani Lischinski, Ohad Fried, and Xi Yin. 2022a · 2022
Closest in time.
Text2LIVE: Text-Driven Layered Image and Video Editing
Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. 2022 · 2022
Closest in time.
VQGAN-CLIP: Open Domain Image Generation and Editing with Natural Language Guidance
Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff. 2022 · 2022
Closest in time.
Make-a-scene: Scene-based text-to-image generation with human priors
Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. 2022 · 2022
Closest in time.
StyleGAN-NADA: CLIP-guided domain adaptation of image generators
Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022 · 2022
Closest in time.
Prompt-to-prompt image editing with cross attention control
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2022 · 2022
Closest in time.
Repaint: Inpainting using denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11461–11471
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022 · 2022
Closest in time.
No Token Left Behind: Explainability-Aided Image Classification and Generation
Roni Paiss, Hila Chefer, and Lior Wolf. 2022 · 2022
Closest in time.
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Closest in time.
Palette: Image-to-image diffusion models. In ACM SIGGRAPH 2022 Conference Proceedings . 1–10
Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. 2022a · 2022
Closest in time.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al · 2022
Closest in time.
Stitch it in Time: GAN-Based Facial Editing of Real Videos
Rotem Tzaban, Ron Mokady, Rinon Gal, Amit H Bermano, and Daniel Cohen-Or. 2022 · 2022
Closest in time.
CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP
Zihao Wang, Wei Liu, Qian He, Xinglong Wu, and Zili Yi. 2022 · 2022
Closest in time.
Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al · 2022
Closest in time.
StyleCLIP: Text-driven manipulation of StyleGAN imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 2085–2094
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021 · 2094
Closest in time.