Fetching the paper…
Reading the bibliography…
Current large-scale generative models have impressive efficiency in generating high-quality images based on text prompts.
“Microsoft coco: Common objects in context,”
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick, · 2014
Earlier work this paper cites.
“U-net: Convolutional networks for biomedical image segmentation,”
Olaf Ronneberger, Philipp Fischer, and Thomas Brox, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Gans trained by a two time-scale update rule converge to a local nash equilibrium,”
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter, · 2017
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Earlier work this paper cites.
“Diffusion models beat gans on image synthesis,”
Prafulla Dhariwal and Alexander Nichol, · 2021
Earlier work this paper cites.
“Denoising diffusion implicit models,”
Jiaming Song, Chenlin Meng, and Stefano Ermon, · 2021
Earlier work this paper cites.
“Improved denoising diffusion probabilistic models,”
Alexander Quinn Nichol and Prafulla Dhariwal, · 2021
Cited alongside, same era.
“Benchmark for compositional text-to-image synthesis,”
Dong Huk Park, Samaneh Azadi, Xihui Liu, Trevor Darrell, and Anna Rohrbach, · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Cited alongside, same era.
“You only learn one representation: Unified network for multiple tasks,”
Chien-Yao Wang, I-Hau Yeh, and Hong-Yuan Mark Liao, · 2021
Cited alongside, same era.
“High-resolution image synthesis with latent diffusion models,”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, · 2022
Cited alongside, same era.
“Glide: Towards photorealistic image generation and editing with text-guided diffusion models,”
Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mcgrew, Ilya Sutskever, and Mark Chen, · 2022
Later among the works it cites.
“Compositional visual generation with composable diffusion models,”
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum, · 2022
Later among the works it cites.
“Imagic: Text-based real image editing with diffusion models,”
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani, · 2022
Later among the works it cites.
“Training-free structured diffusion guidance for compositional text-to-image synthesis,”
Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Arjun Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang, · 2022
Later among the works it cites.
“Prompt-to-prompt image editing with cross attention control,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen, · 2022
Cited alongside, same era.
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or, · 2022
Later among the works it cites.
“ediffi: Text-to-image diffusion models with an ensemble of expert denoisers,”
Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al., · 2022
Later among the works it cites.