Fetching the paper…
Reading the bibliography…
The rapid advancement of generative AI has democratized access to powerful tools such as Text-to-Image models.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Earlier work this paper cites.
Slic: Self-supervised learning with iterative clustering for human action videos
Salar Hosseini Khorasgani, Yuxuan Chen, and Florian Shkurti. 2022 · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. 2022 · 2022
Earlier work this paper cites.
Bbtv2: Towards a gradient-free future with large language models
Tianxiang Sun, Zhengfu He, Hong Qian, Yunhua Zhou, Xuan-Jing Huang, and Xipeng Qiu. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
BeautifulPrompt: Towards automatic prompt engineering for text-to-image synthesis
Tingfeng Cao, Chengyu Wang, Bingyan Liu, Ziheng Wu, Jinhui Zhu, and Jun Huang. 2023 · 2023
Earlier work this paper cites.
Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models
Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. 2023 · 2023
Earlier work this paper cites.
Visual programming for step-by-step text-to-image generation and evaluation
Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023 · 2023
Earlier work this paper cites.
Diffusion self-guidance for controllable image generation
Dave Epstein, Allan Jabri, Ben Poole, Alexei Efros, and Aleksander Holynski. 2023 · 2023
Earlier work this paper cites.
Promptmagician: Interactive prompt engineering for text-to-image creation
Yingchaojie Feng, Xingbo Wang, Kam Kwai Wong, Sijia Wang, Yuhong Lu, Minfeng Zhu, Baicheng Wang, and Wei Chen. 2023 · 2023
Earlier work this paper cites.
Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini, and Zeynep Akata. 2023 · 2023
Earlier work this paper cites.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. 2023 · 2023
Cited alongside, same era.
Aligning text-to-image models using human feedback
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu. 2023 · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023 · 2023
Cited alongside, same era.
Grips: Gradient-free, edit-based instruction search for prompting large language models
Is it ai or is it me? understanding users’ prompt journey with text-to-image generative ai tools
Atefeh Mahdavi Goloujeh, Anne Sullivan, and Brian Magerko. 2024 · 2024
Later among the works it cites.
Midjourney
Midjourney. 2024 · 2024
Later among the works it cites.
Dynamic prompt optimizing for text-to-image generation
Wenyi Mo, Tianyu Zhang, Yalong Bai, Bing Su, Ji-Rong Wen, and Qing Yang. 2024 · 2024
Later among the works it cites.
Dreamsync: Aligning text-to-image generation with image understanding feedback
Jiao Sun, Deqing Fu, Yushi Hu, Su Wang, Royi Rassin, Da-Cheng Juan, Dana Alon, Charles Herrmann, Sjoerd van Steenkiste, Ranjay Krishna, et al · 2024
Later among the works it cites.
Diffusion model alignment using direct preference optimization
Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal. 2023 · 2023
Cited alongside, same era.
Automatic prompt optimization with" gradient descent" and beam search
Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng. 2023 · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Cited alongside, same era.
Toward human readable prompt tuning: Kubrick’s the shining is a good movie, and a good prompt too?
Weijia Shi, Xiaochuang Han, Hila Gonen, Ari Holtzman, Yulia Tsvetkov, and Luke Zettlemoyer · 2023
Cited alongside, same era.
Text-to-image diffusion models in generative ai: A survey
Chenshuang Zhang, Chaoning Zhang, Mengchun Zhang, and In So Kweon. 2023 · 2023
Cited alongside, same era.
Promptbreeder: Self-referential self-improvement via prompt evolution
Chrisantha Fernando, Dylan Sunil Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. 2024 · 2024
Cited alongside, same era.
Large language model based multi-agents: A survey of progress and challenges
Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang. 2024 · 2024
Cited alongside, same era.
Localized zeroth-order prompt optimization
Wenyang Hu, Yao Shu, Zongmin Yu, Zhaoxuan Wu, Xiaoqiang Lin, Zhongxiang Dai, See-Kiong Ng, and Bryan Kian Hsiang Low. 2024 · 2024
Cited alongside, same era.
Xidong Wu, Sumin Jo, Yiming Zeng, Arun Das, Ting-He Zhang, Parth Patel, Yuanjing Wei, Lei Li, Shou-Jiang Gao, Jianqiu Zhang, et al. 2024 · 2024
Later among the works it cites.
Revisiting opro: The limitations of small-scale llms as optimizers
Tuo Zhang, Jinyue Yuan, and Salman Avestimehr. 2024 · 2024
Later among the works it cites.
Can we generate images with cot? let’s verify and reinforce image generation step by step
Ziyu Guo, Renrui Zhang, Chengzhuo Tong, Zhizheng Zhao, Peng Gao, Hongsheng Li, and Pheng-Ann Heng. 2025 · 2025
Closest in time.
From llm-anation to llm-orchestrator: Coordinating small models for data labeling
Yao Lu, Zhaiyuan Ji, Jiawei Du, Yu Shanqing, Qi Xuan, and Tianyi Zhou. 2025 · 2025
Closest in time.
Hrft: Mining high-frequency risk factor collections end-to-end via transformer
Wenyan Xu, Rundong Wang, Chen Li, Yonghong Hu, and Zhonghua Lu. 2025a · 2025
Closest in time.
Time-llama: Adapting large language models for time series modeling via dynamic low-rank adaptation
J. Zhang, J. Gao, W. Ouyang, W. Zhu, and H.Y. Leong. 2025a · 2025
Closest in time.
Styleclip: Text-driven manipulation of stylegan imagery
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021 · 2094
Closest in time.
Human preference score: Better aligning text-to-image models with human preference
Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, and Hongsheng Li. 2023b · 2096
Closest in time.