Fetching the paper…
Reading the bibliography…
We introduce a new architecture for personalization of text-to-image diffusion models, coined Mixture-of-Attention (MoA).
Multi-concept customization of text-to-image diffusion. In CVPR . 1931–1941
Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. 2023a · 1941
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991 · 1991
Earlier work this paper cites.
Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision . 3730–3738
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015 · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering. In CVPR . 815–823
Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015 · 2015
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In International Conference on Learning Representations
Noam Shazeer, *Azalia Mirhoseini, *Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks. In CVPR . 4401–4410
Tero Karras, Samuli Laine, and Timo Aila. 2019 · 2019
Earlier work this paper cites.
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and P. Abbeel. 2020 · 2020
Earlier work this paper cites.
Denoising Diffusion Implicit Models
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020a · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020b · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. 2021 · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models. In International Conference on Machine Learning . PMLR, 8162–8171
Alexander Quinn Nichol and Prafulla Dhariwal. 2021 · 2021
Earlier work this paper cites.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Jason Weston, et al · 2021
Earlier work this paper cites.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2021 · 2021
Earlier work this paper cites.
Low-rank Adaptation for Fast Text-to-Image Diffusion Fine-tuning
2022 · 2022
Earlier work this paper cites.
Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1290–1299
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. 2022 · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2022 · 2022
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022 · 2022
Earlier work this paper cites.
Accelerate: Training and inference at scale made simple, efficient and adaptable
Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan. 2022 · 2022
Earlier work this paper cites.
Classifier-Free Diffusion Guidance
Jonathan Ho. 2022 · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models. In ICLR
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
DiffuseVAE: Efficient, Controllable and High-Fidelity Generation from Low-Dimensional Latents
Kushagra Pandey, Avideep Mukherjee, Piyush Rai, and Abhishek Kumar. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10684–10695
Photomaker: Customizing realistic human photos via stacked id embedding
Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming-Ming Cheng, and Ying Shan. 2023a · 2023
Later among the works it cites.
Cones 2: Customizable image synthesis with multiple subjects
Zhiheng Liu, Yifei Zhang, Yujun Shen, Kecheng Zheng, Kai Zhu, Ruili Feng, Yu Liu, Deli Zhao, Jingren Zhou, and Yang Cao. 2023a · 2023
Later among the works it cites.
Cones 2: Customizable image synthesis with multiple subjects
Zhiheng Liu, Yifei Zhang, Yujun Shen, Kecheng Zheng, Kai Zhu, Ruili Feng, Yu Liu, Deli Zhao, Jingren Zhou, and Yang Cao. 2023b · 2023
Later among the works it cites.
Null-text inversion for editing real images using guided diffusion models. In CVPR . 6038–6047
Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding. In NeurIPS . 36479–36494
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Cited alongside, same era.
Unipc: A unified predictor-corrector framework for fast sampling of diffusion models
Wenliang Zhao, Lujia Bai, Yongming Rao, Jie Zhou, and Jiwen Lu. 2023 · 2022
Cited alongside, same era.
A Neural Space-Time Representation for Text-to-Image Personalization
Yuval Alaluf, Elad Richardson, Gal Metzer, and Daniel Cohen-Or. 2023 · 2023
Cited alongside, same era.
Domain-agnostic tuning-encoder for fast personalization of text-to-image models. In SIGGRAPH Asia 2023 Conference Papers . 1–10
Moab Arar, Rinon Gal, Yuval Atzmon, Gal Chechik, Daniel Cohen-Or, Ariel Shamir, and Amit H. Bermano. 2023 · 2023
Cited alongside, same era.
Break-A-Scene: Extracting Multiple Concepts from a Single Image
Omri Avrahami, Kfir Aberman, Ohad Fried, Daniel Cohen-Or, and Dani Lischinski. 2023a · 2023
Cited alongside, same era.
The Chosen One: Consistent Characters in Text-to-Image Diffusion Models
Omri Avrahami, Amir Hertz, Yael Vinker, Moab Arar, Shlomi Fruchter, Ohad Fried, Daniel Cohen-Or, and Dani Lischinski. 2023b · 2023
Cited alongside, same era.
MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation
Omer Bar-Tal, Lior Yariv, Yaron Lipman, and Tali Dekel. 2023 · 2023
Cited alongside, same era.
Ryan Po, Guandao Yang, Kfir Aberman, and Gordon Wetzstein. 2023 · 2023
Later among the works it cites.
Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Wei Wei, Tingbo Hou, Yael Pritch, Neal Wadhwa, Michael Rubinstein, and Kfir Aberman. 2023b · 2023
Later among the works it cites.
Instantbooth: Personalized text-to-image generation without test-time finetuning
Jing Shi, Wei Xiong, Zhe Lin, and Hyun Joon Jung. 2023 · 2023
Later among the works it cites.
Key-locked rank one editing for text-to-image personalization. In ACM SIGGRAPH 2023 Conference Proceedings . 1–11
Yoad Tewel, Rinon Gal, Gal Chechik, and Yuval Atzmon. 2023 · 2023
Later among the works it cites.
P + P+ : Extended Textual Conditioning in Text-to-Image Generation
Andrey Voynov, Qinghao Chu, Daniel Cohen-Or, and Kfir Aberman. 2023 · 2023
Later among the works it cites.
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation
Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. 2023 · 2023
Later among the works it cites.
FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention
Guangxuan Xiao, Tianwei Yin, William T Freeman, Frédo Durand, and Song Han. 2023 · 2023
Later among the works it cites.
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023 · 2023
Later among the works it cites.
Prospect: Prompt spectrum for attribute-aware personalization of diffusion models
Yuxin Zhang, Weiming Dong, Fan Tang, Nisha Huang, Haibin Huang, Chongyang Ma, Tong-Yee Lee, Oliver Deussen, and Changsheng Xu. 2023a · 2023
Later among the works it cites.
LCM-Lookahead for Encoder-based Text-to-Image Personalization
Rinon Gal, Or Lichter, Elad Richardson, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2024 · 2024
Closest in time.
Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models
Yuchao Gu, Xintao Wang, Jay Zhangjie Wu, Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao, Shuning Chang, Weijia Wu, et al · 2024
Closest in time.
Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing
Dongxu Li, Junnan Li, and Steven Hoi. 2024 · 2024
Closest in time.
Training-Free Consistent Text-to-Image Generation
Yoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten, Lior Wolf, Gal Chechik, and Yuval Atzmon. 2024 · 2024
Closest in time.
Instantid: Zero-shot identity-preserving generation in seconds
Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, and Anthony Chen. 2024 · 2024
Closest in time.
PeRFlow: Accelerating Diffusion models via Piecewise Rectified Flow
Hanshu Yan, Xingchao Liu, Jiachun Pan, Jun Hao Liew, Qiang Liu, and Jiashi Feng. 2024 · 2024
Closest in time.