Fetching the paper…
Reading the bibliography…
Training-free diffusion models have achieved remarkable progress in generating multi-subject consistent images within open-domain scenarios.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Wasserstein generative adversarial networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Earlier work this paper cites.
Improved training of Wasserstein GANs
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Earlier work this paper cites.
Imagine this! scripts to compositions to videos
Tanmay Gupta, Dustin Schwenk, Ali Farhadi, Derek Hoiem, and Aniruddha Kembhavi · 2018
Earlier work this paper cites.
Storytelling and visualization: An extended survey
Chao Tong, Richard Roberts, Rita Borgo, Sean Walton, Robert S Laramee, Kodzo Wegba, Aidong Lu, Yun Wang, Huamin Qu, Qiong Luo, et al · 2018
Earlier work this paper cites.
StoryGAN: A sequential conditional gan for story visualization
Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
How much position information do convolutional neural networks encode?
Md Amirul Islam, Sen Jia, and Neil DB Bruce · 2020
Earlier work this paper cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2021
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Earlier work this paper cites.
Integrating visuospatial, linguistic, and commonsense structure into story visualization
Adyasha Maharana and Mohit Bansal · 2021
Earlier work this paper cites.
Improving generation and evaluation of visual stories via semantic consistency
Adyasha Maharana, Darryl Hannan, and Mohit Bansal · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Character-centric story visualization via visual planning and token alignment
Hong Chen, Rujun Han, Te-Lin Wu, Hideki Nakayama, and Nanyun Peng · 2022
Earlier work this paper cites.
BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Earlier work this paper cites.
DPM-solver++: Fast solver for guided sampling of diffusion probabilistic models
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu · 2022
Earlier work this paper cites.
AI illustrator: Translating raw descriptions into images by prompt-based cross-modal generation
Yiyang Ma, Huan Yang, Bei Liu, Jianlong Fu, and Jiaying Liu · 2022
Earlier work this paper cites.
StoryDALL-E: Adapting pretrained text-to-image transformers for story continuation
Adyasha Maharana, Darryl Hannan, and Mohit Bansal · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Cited alongside, same era.
Domain-agnostic tuning-encoder for fast personalization of text-to-image models
Moab Arar, Rinon Gal, Yuval Atzmon, Gal Chechik, Daniel Cohen-Or, Ariel Shamir, and Amit H. Bermano · 2023
Cited alongside, same era.
MasaCtrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng · 2023
Cited alongside, same era.
TeViS: Translating text synopses to video storyboards
Xu Gu, Yuchong Sun, Feiyue Ni, Shizhe Chen, Xihua Wang, Ruihua Song, Boyuan Li, and Xiang Cao · 2023
Cited alongside, same era.
Learning profitable NFT image diffusions via multiple visual-policy guided reinforcement learning
Huiguo He, Tianfu Wang, Huan Yang, Jianlong Fu, Nicholas Jing Yuan, Jian Yin, Hongyang Chao, and Qi Zhang · 2023
Position, padding and predictions: A deeper look at position information in cnns
Md Amirul Islam, Matthew Kowal, Sen Jia, Konstantinos G Derpanis, and Neil DB Bruce · 2024
Closest in time.
Identity decoupling for multi-subject personalization of text-to-image models, 2024
Sangwon Jang, Jaehyeong Jo, Kimin Lee, and Sung Ju Hwang · 2024
Closest in time.
InstantFamily: Masked attention for zero-shot multi-id image generation
Chanran Kim, Jeongin Lee, Shichang Joung, Bongmo Kim, and Yeul-Min Baek · 2024
Closest in time.
OMG: Occlusion-friendly personalized multi-concept generation in diffusion models
Zhe Kong, Yong Zhang, Tianyu Yang, Tao Wang, Kaihao Zhang, Bizhu Wu, Guanying Chen, Wei Liu, and Wenhan Luo · 2024
Closest in time.
Direct consistency optimization for compositional text-to-image personalization
Kyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, and Jinwoo Shin · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Cited alongside, same era.
Unveiling the mask of position-information pattern through the mist of image features
Chieh Hubert Lin, Hung-Yu Tseng, Hsin-Ying Lee, Maneesh Kumar Singh, and Ming-Hsuan Yang · 2023
Cited alongside, same era.
SDXL: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach · 2023
Cited alongside, same era.
Make-a-Story: Visual memory conditioned consistent story generation
Tanzila Rahman, Hsin-Ying Lee, Jian Ren, Sergey Tulyakov, Shweta Mahajan, and Leonid Sigal · 2023
Cited alongside, same era.
DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman · 2023
Cited alongside, same era.
Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation
Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu · 2023
Cited alongside, same era.
IP-Adapter: Text compatible image prompt adapter for text-to-image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang · 2023
Cited alongside, same era.
Closest in time.
Intelligent grimm-open-ended visual storytelling via latent diffusion models
Chang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang, Yanfeng Wang, and Weidi Xie · 2024
Closest in time.
Subject-Diffusion: Open domain personalized text-to-image generation without test-time fine-tuning
Jian Ma, Junhao Liang, Chen Chen, and Haonan Lu · 2024
Closest in time.
Synthesizing coherent story with auto-regressive latent diffusion models
Xichen Pan, Pengda Qin, Yuhong Li, Hui Xue, and Wenhu Chen · 2024
Closest in time.
PortraitBooth: A versatile portrait model for fast identity-preserved personalization
Xu Peng, Junwei Zhu, Boyuan Jiang, Ying Tai, Donghao Luo, Jiangning Zhang, Wei Lin, Taisong Jin, Chengjie Wang, and Rongrong Ji · 2024
Closest in time.
Grounded SAM: Assembling open-world models for diverse visual tasks
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al · 2024
Closest in time.
Image-based video game asset generation and evaluation using deep learning: a systematic review of methods and applications
Rafael Ribeiro, Alexandre Valle de Carvalho, and Nelson Bilber Rodrigues · 2024
Closest in time.
Create your world: Lifelong text-to-image diffusion
Gan Sun, Wenqi Liang, Jiahua Dong, Jun Li, Zhengming Ding, and Yang Cong · 2024
Closest in time.
Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis
Kolors Team · 2024
Closest in time.
Training-free consistent text-to-image generation
Yoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten, Lior Wolf, Gal Chechik, and Yuval Atzmon · 2024
Closest in time.
High-fidelity person-centric subject-to-image synthesis
Yibin Wang, Weizhong Zhang, Jianwei Zheng, and Cheng Jin · 2024
Closest in time.
Jedi: Joint-image diffusion models for finetuning-free personalized text-to-image generation
Yu Zeng, Vishal M Patel, Haochen Wang, Xun Huang, Ting-Chun Wang, Ming-Yu Liu, and Yogesh Balaji · 2024
Closest in time.
StoryMaker: Towards holistic consistent characters in text-to-image generation
Zhengguang Zhou, Jing Li, Huaxia Li, Nemo Chen, and Xu Tang · 2024
Closest in time.
MultiBooth: Towards generating all your concepts in an image from text
Chenyang Zhu, Kai Li, Yue Ma, Chunming He, and Li Xiu · 2024
Closest in time.
One-Prompt-One-Story: Free-lunch consistent text-to-image generation using a single prompt
Tao Liu, Kai Wang, Senmao Li, Joost van de Weijer, Fahad Shahbaz Khan, Shiqi Yang, Yaxing Wang, Jian Yang, and Ming-Ming Cheng · 2025
Closest in time.
StoryDiffusion: Consistent self-attention for long-range image and video generation
Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou · 2025
Closest in time.