Fetching the paper…
Reading the bibliography…
Current image generation systems produce high-quality images but struggle with ambiguous user prompts, making interpretation of actual user intentions difficult.
“Diverse beam search: Decoding diverse solutions from neural sequence models,”
Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra, · 2016
Earlier work this paper cites.
“Proximal policy optimization algorithms,”
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov, · 2017
Earlier work this paper cites.
“Gans trained by a two time-scale update rule converge to a local nash equilibrium,”
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter, · 2017
Earlier work this paper cites.
“The unreasonable effectiveness of deep features as a perceptual metric,”
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang, · 2018
Earlier work this paper cites.
“Retrieval-augmented generation for knowledge-intensive nlp tasks,”
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al., · 2020
Earlier work this paper cites.
“Adversarial text-to-image synthesis: A review,”
Stanislav Frolov, Tobias Hinz, Federico Raue, Jörn Hees, and Andreas Dengel, · 2021
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Earlier work this paper cites.
“Glide: Towards photorealistic image generation and editing with text-guided diffusion models,”
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen, · 2021
Earlier work this paper cites.
“Photorealistic text-to-image diffusion models with deep language understanding,” 2022
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi, · 2022
Earlier work this paper cites.
“High-resolution image synthesis with latent diffusion models,”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer, · 2022
Earlier work this paper cites.
“First contact: Unsupervised human-machine co-adaptation via mutual information maximization,”
Siddharth Reddy, Sergey Levine, and Anca Dragan, · 2022
Cited alongside, same era.
“Large language models can self-improve,”
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han, · 2022
Cited alongside, same era.
“Prompt-to-prompt image editing with cross attention control,”
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or, · 2022
Cited alongside, same era.
“Make-a-scene: Scene-based text-to-image generation with human priors,”
Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman, · 2022
Cited alongside, same era.
“Improving image generation with better captions,”
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al., · 2023
“Resolving ambiguities in text-to-image generative models,”
Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard Zemel, Aram Galstyan, and Rahul Gupta, · 2023
Later among the works it cites.
“Rich human feedback for text-to-image generation,”
Youwei Liang, Junfeng He, Gang Li, Peizhao Li, Arseniy Klimovskiy, Nicholas Carolan, Jiao Sun, Jordi Pont-Tuset, Sarah Young, Feng Yang, et al., · 2023
Later among the works it cites.
“Better aligning text-to-image models with human preference,”
Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, and Hongsheng Li, · 2023
Later among the works it cites.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang, · 2023
Later among the works it cites.
“Blended latent diffusion,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Chatgpt is not all you need. a state of the art review of large generative ai models,”
Roberto Gozalo-Brizuela and Eduardo C Garrido-Merchan, · 2023
Cited alongside, same era.
“Diffusion models in vision: A survey,”
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah, · 2023
Cited alongside, same era.
“Matryoshka diffusion models,” 2023
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Josh Susskind, and Navdeep Jaitly, · 2023
Cited alongside, same era.
“A survey on video diffusion models,”
Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, and Yu-Gang Jiang, · 2023
Cited alongside, same era.
“Imagen editor and editbench: Advancing and evaluating text-guided image inpainting,”
S. Wang, C. Saharia, C. Montgomery, J. Pont-Tuset, S. Noy, S. Pellegrini, Y. Onoe, S. Laszlo, D. J. Fleet, R. Soricut, J. Baldridge, M. Norouzi, P. Anderson, and W. Chan, · 2023
Cited alongside, same era.
Omri Avrahami, Ohad Fried, and Dani Lischinski, · 2023
Later among the works it cites.
“Instructpix2pix: Learning to follow image editing instructions,”
Tim Brooks, Aleksander Holynski, and Alexei A Efros, · 2023
Later among the works it cites.
“Aligning text-to-image models using human feedback,”
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu, · 2023
Later among the works it cites.
“Scaling rectified flow transformers for high-resolution image synthesis,” 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach, · 2024
Later among the works it cites.
“Cogview3: Finer and faster text-to-image generation via relay diffusion,” 2024
Wendi Zheng, Jiayan Teng, Zhuoyi Yang, Weihan Wang, Jidong Chen, Xiaotao Gu, Yuxiao Dong, Ming Ding, and Jie Tang, · 2024
Later among the works it cites.
“Imagereward: Learning and evaluating human preferences for text-to-image generation,”
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong, · 2024
Later among the works it cites.