Fetching the paper…
Reading the bibliography…
Fine-tuning facilitates the adaptation of text-to-image generative models to novel concepts (e.g., styles and portraits), empowering users to forge creatively customized content.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Color harmonization
Daniel Cohen-Or, Olga Sorkine, Ran Gal, Tommer Leyvand, and Ying-Qing Xu. 2006 · 2006
Earlier work this paper cites.
AVA: A large-scale database for aesthetic visual analysis. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2408–2415
Naila Murray, Luca Marchesotti, and Florent Perronnin. 2012 · 2012
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. 2016 · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Earlier work this paper cites.
Visual question answering: A survey of methods and datasets
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. 2017 · 2017
Earlier work this paper cites.
Gradio: Hassle-free sharing and testing of ml models in the wild
Abubakar Abid, Ali Abdalla, Ali Abid, Dawood Khan, Abdulrahman Alfozan, and James Zou. 2019 · 2019
Earlier work this paper cites.
Salient object detection: A survey
Ali Borji, Ming-Ming Cheng, Qibin Hou, Huaizu Jiang, and Jia Li. 2019 · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4401–4410
Tero Karras, Samuli Laine, and Timo Aila. 2019 · 2019
Earlier work this paper cites.
The emergence of deepfake technology: A review
Mika Westerlund. 2019 · 2019
Earlier work this paper cites.
8-bit optimizers via block-wise quantization
Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International Conference on Machine Learning . 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation. In International Conference on Machine Learning . 8821–8831
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Earlier work this paper cites.
Kohya’s Stable Diffusion trainers
2022 · 2022
Earlier work this paper cites.
LibLibAI
2022 · 2022
Earlier work this paper cites.
Stable Diffusion Web UI
2022 · 2022
Earlier work this paper cites.
TISE: Bag of metrics for text-to-image synthesis evaluation. In European Conference on Computer Vision . Springer, 594–609
Tan M Dinh, Rang Nguyen, and Binh-Son Hua. 2022 · 2022
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022a · 2022
Earlier work this paper cites.
StyleGAN-NADA: CLIP-guided domain adaptation of image generators
Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022b · 2022
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
We-toon: A Communication Support System between Writers and Artists in Collaborative Webtoon Sketch Revision. In Proceedings of the Annual ACM Symposium on User Interface Software and Technology . 1–14
Hyung-Kwon Ko, Subin An, Gwanmo Park, Seung Kwon Kim, Daesik Kim, Bohyoung Kim, Jaemin Jo, and Jinwook Seo. 2022 · 2022
Cited alongside, same era.
Design guidelines for prompt engineering text-to-image generative models. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–23
Vivian Liu and Lydia B Chilton. 2022 · 2022
Cited alongside, same era.
Opal: Multimodal image generation for news illustration. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology . 1–17
Vivian Liu, Han Qiao, and Lydia Chilton. 2022 · 2022
Cited alongside, same era.
Repaint: Inpainting using denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11461–11471
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022 · 2022
Cited alongside, same era.
Semantic-SAM: Segment and Recognize Anything at Any Granularity
Feng Li, Hao Zhang, Peize Sun, Xueyan Zou, Shilong Liu, Jianwei Yang, Chunyuan Li, Lei Zhang, and Jianfeng Gao. 2023b · 2023
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023a · 2023
Later among the works it cites.
Taskmatrix. AI: Completing tasks by connecting foundation models with millions of apis
Yaobo Liang, Chenfei Wu, Ting Song, Wenshan Wu, Yan Xia, Yu Liu, Yang Ou, Shuai Lu, Lei Ji, Shaoguang Mao, et al · 2023
Later among the works it cites.
Application potential of stable diffusion in different stages of industrial design. In International Conference on Human-Computer Interaction . Springer, 590–609
Miao Liu and Yifei Hu. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rethinking the Role of Demonstrations: What makes In-context Learning Work?. In EMNLP
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al · 2022
Cited alongside, same era.
Adobe Firefly
2023 · 2023
Cited alongside, same era.
Midjourney
2023 · 2023
Cited alongside, same era.
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al · 2023
Later among the works it cites.
Beyond Text-to-Image: Multimodal Prompts to Explore Generative AI. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–6
Vivian Liu. 2023 · 2023
Later among the works it cites.
The programmer’s assistant: Conversational interaction with a large language model for software development. In Proceedings of the 28th International Conference on Intelligent User Interfaces . 491–514
Steven I Ross, Fernando Martinez, Stephanie Houde, Michael Muller, and Justin D Weisz. 2023 · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22500–22510
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023 · 2023
Later among the works it cites.
Glaze: Protecting artists from style mimicry by text-to-image models
Shawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng, Rana Hanocka, and Ben Y Zhao. 2023 · 2023
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023 · 2023
Later among the works it cites.
Diffusion art or digital forgery? investigating data replication in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6048–6058
Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Collaborative Diffusion: Boosting Designerly Co-Creation with Generative AI. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–8
Mathias Peter Verheijden and Mathias Funk. 2023 · 2023
Later among the works it cites.
Concept Decomposition for Visual Exploration and Inspiration
Yael Vinker, Andrey Voynov, Daniel Cohen-Or, and Ariel Shamir. 2023 · 2023
Later among the works it cites.
Better Aligning Text-to-Image Models with Human Preference
Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, and Hongsheng Li. 2023a · 2023
Later among the works it cites.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. 2023 · 2023
Later among the works it cites.
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
Jingfeng Yao, Xinggang Wang, Lang Ye, and Wenyu Liu. 2023 · 2023
Later among the works it cites.
Inpaint anything: Segment anything meets image inpainting
Tao Yu, Runseng Feng, Ruoyu Feng, Jinming Liu, Xin Jin, Wenjun Zeng, and Zhibo Chen. 2023 · 2023
Later among the works it cites.
A Comprehensive Survey on Segment Anything Model for Vision and Beyond
Chunhui Zhang, Li Liu, Yawen Cui, Guanjie Huang, Weilin Lin, Yiqian Yang, and Yuehong Hu. 2023b · 2023
Later among the works it cites.
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable and Controllable Text-Guided Face Manipulation. In ACM SIGGRAPH 2023 Conference Proceedings . 1–9
Chenliang Zhou, Fangcheng Zhong, and Cengiz Öztireli. 2023 · 2023
Later among the works it cites.