Fetching the paper…
Reading the bibliography…
Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging.
Deep sparse rectifier neural networks. In Proceedings of the fourteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 315–323
Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011 · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma. 2013 · 2013
Earlier work this paper cites.
Decoupled weight decay regularization
I Loshchilov. 2017 · 2017
Earlier work this paper cites.
Gradio: Hassle-Free Sharing and Testing of ML Models in the Wild
Abubakar Abid, Ali Abdalla, Ali Abid, Dawood Khan, Abdulrahman Alfozan, and James Zou. 2019 · 2019
Earlier work this paper cites.
A benchmark and baseline for language-driven image editing. In Proceedings of the Asian Conference on Computer Vision
Jing Shi, Ning Xu, Trung Bui, Franck Dernoncourt, Zheng Wen, and Chenliang Xu. 2020 · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Earlier work this paper cites.
Resolution-robust large mask inpainting with fourier convolutions. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 2149–2159
Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. 2022 · 2022
Earlier work this paper cites.
Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18392–18402
Tim Brooks, Aleksander Holynski, and Alexei A Efros. 2023 · 2023
Earlier work this paper cites.
Pixart- α \alpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al · 2023
Earlier work this paper cites.
Eva: Exploring the limits of masked visual representation learning at scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 19358–19369
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. 2023 · 2023
Earlier work this paper cites.
Guiding instruction-based image editing via multimodal large language models
Tsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang, Yinfei Yang, and Zhe Gan. 2023 · 2023
Earlier work this paper cites.
Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4015–4026
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Earlier work this paper cites.
Dreamedit: Subject-driven image editing
Tianle Li, Max Ku, Cong Wei, and Wenhu Chen. 2023a · 2023
Earlier work this paper cites.
Kosmos-g: Generating images in context with multimodal large language models
Xichen Pan, Li Dong, Shaohan Huang, Zhiliang Peng, Wenhu Chen, and Furu Wei. 2023 · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4195–4205
William Peebles and Saining Xie. 2023 · 2023
Earlier work this paper cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 22500–22510
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023 · 2023
Earlier work this paper cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms
MosaicML NLP Team et al · 2023
Earlier work this paper cites.
Paint by example: Exemplar-based image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18381–18391
Binxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang, Xuejin Chen, Xiaoyan Sun, Dong Chen, and Fang Wen. 2023 · 2023
Earlier work this paper cites.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023 · 2023
Earlier work this paper cites.
Inst-inpaint: Instructing to remove objects with diffusion models
Ahmet Burak Yildirim, Vedat Baday, Erkut Erdem, Aykut Erdem, and Aysegul Dundar. 2023 · 2023
Earlier work this paper cites.
Zero-shot image editing with reference imitation
Xi Chen, Yutong Feng, Mengting Chen, Yiyang Wang, Shilong Zhang, Yu Liu, Yujun Shen, and Hengshuang Zhao. 2024a · 2024
Earlier work this paper cites.
UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics
Xi Chen, Zhifei Zhang, He Zhang, Yuqian Zhou, Soo Ye Kim, Qing Liu, Yijun Li, Jianming Zhang, Nanxuan Zhao, Yilin Wang, et al · 2024
Earlier work this paper cites.
Improving Diffusion Models for Authentic Virtual Try-on in the Wild
Yisol Choi, Sangkyung Kwak, Kyungmin Lee, Hyungwon Choi, and Jinwoo Shin. 2024 · 2024
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first International Conference on Machine Learning
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Cited alongside, same era.
SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
Yuying Ge, Sijie Zhao, Chen Li, Yixiao Ge, and Ying Shan. 2024 · 2024
Cited alongside, same era.
ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer
Ominicontrol: Minimal and universal control for diffusion transformer
Zhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue, and Xinchao Wang. 2024 · 2024
Later among the works it cites.
Do We Need to Design Specific Diffusion Models for Different Tasks? Try ONE-PIC
Ming Tao, Bing-Kun Bao, Yaowei Wang, and Changsheng Xu. 2024 · 2024
Later among the works it cites.
Qwen2.5: A Party of Foundation Models
Qwen Team. 2024 · 2024
Later among the works it cites.
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhen Han, Zeyinzi Jiang, Yulin Pan, Jingfeng Zhang, Chaojie Mao, Chenwei Xie, Yu Liu, and Jingren Zhou. 2024 · 2024
Cited alongside, same era.
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion
Jixuan He, Wanhua Li, Ye Liu, Junsik Kim, Donglai Wei, and Hanspeter Pfister. 2024a · 2024
Cited alongside, same era.
Freeedit: Mask-free reference-based image editing with multi-modal instruction
Runze He, Kai Ma, Linjiang Huang, Shaofei Huang, Jialin Gao, Xiaoming Wei, Jiao Dai, Jizhong Han, and Si Liu. 2024b · 2024
Cited alongside, same era.
Instruct-Imagen: Image generation with multi-modal instruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4754–4763
Hexiang Hu, Kelvin CK Chan, Yu-Chuan Su, Wenhu Chen, Yandong Li, Kihyuk Sohn, Yang Zhao, Xue Ben, Boqing Gong, William Cohen, et al · 2024
Cited alongside, same era.
Group diffusion transformers are unsupervised multitask learners
Lianghua Huang, Wei Wang, Zhi-Fan Wu, Huanzhang Dou, Yupeng Shi, Yutong Feng, Chen Liang, Yu Liu, and Jingren Zhou. 2024a · 2024
Cited alongside, same era.
Hq-edit: A high-quality dataset for instruction-based image editing
Mude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi, Heng Wang, Peng Wang, Yuyin Zhou, and Cihang Xie. 2024 · 2024
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Cited alongside, same era.
Learning to Customize Text-to-Image Diffusion In Diverse Context
Taewook Kim, Wei Chen, and Qiang Qiu. 2024 · 2024
Cited alongside, same era.
Shitao Xiao, Yueze Wang, Junjie Zhou, Huaying Yuan, Xingrun Xing, Ruiran Yan, Shuting Wang, Tiejun Huang, and Zheng Liu. 2024 · 2024
Later among the works it cites.
InsightEdit: Towards Better Instruction Following for Image Editing
Yingjing Xu, Jie Kong, Jiazhi Wang, Xiao Pan, Bo Lin, and Qiang Liu. 2024 · 2024
Later among the works it cites.
xgen-mm (blip-3): A family of open large multimodal models
Le Xue, Manli Shu, Anas Awadalla, Jun Wang, An Yan, Senthil Purushwalkam, Honglu Zhou, Viraj Prabhu, Yutong Dai, Michael S Ryoo, et al · 2024
Later among the works it cites.
DreamMix: Decoupling Object Attributes for Enhanced Editability in Customized Image Inpainting
Yicheng Yang, Pengxiang Li, Lu Zhang, Liqian Ma, Ping Hu, Siyu Du, Yunzhi Zhuge, Xu Jia, and Huchuan Lu. 2024 · 2024
Later among the works it cites.
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
Qifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan, Yang Wu, Xiaoyang Wan, Juncheng Li, Siliang Tang, Hanwang Zhang, and Yueting Zhuang. 2024 · 2024
Later among the works it cites.
Magicbrush: A manually annotated dataset for instruction-guided image editing
Kai Zhang, Lingbo Mo, Wenhu Chen, Huan Sun, and Yu Su. 2024 · 2024
Later among the works it cites.
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
Haozhe Zhao, Xiaojian Ma, Liang Chen, Shuzheng Si, Rujie Wu, Kaikai An, Peiyu Yu, Minjia Zhang, Qing Li, and Baobao Chang. 2024 · 2024
Later among the works it cites.
EraseDraw: Learning to Insert Objects by Erasing Them from Images. In European Conference on Computer Vision . Springer, 144–160
Alper Canberk, Maksym Bondarenko, Ege Ozguroglu, Ruoshi Liu, and Carl Vondrick. 2025 · 2025
Closest in time.
Teng-Fang Hsiao, Bo-Kai Ruan, Yi-Lun Wu, Tzu-Ling Lin, and Hong-Han Shuai. 2025 · 2025
Closest in time.
WeGen: A Unified Model for Interactive Multimodal Generation as We Chat
Zhipeng Huang, Shaobin Zhuang, Canmiao Fu, Binxin Yang, Ying Zhang, Chong Sun, Zhizheng Zhang, Yali Wang, Chen Li, and Zheng-Jun Zha. 2025 · 2025
Closest in time.
Flux Already Knows–Activating Subject-Driven Image Generation without Training
Hao Kang, Stathi Fotiadis, Liming Jiang, Qing Yan, Yumin Jia, Zichuan Liu, Min Jin Chong, and Xin Lu. 2025 · 2025
Closest in time.
Blobctrl: A unified and flexible framework for element-level image generation and editing
Yaowei Li, Lingen Li, Zhaoyang Zhang, Xiaoyu Li, Guangzhi Wang, Hongxiang Li, Xiaodong Cun, Ying Shan, and Yuexian Zou. 2025b · 2025
Closest in time.
VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning
Zhong-Yu Li, Ruoyi Du, Juncheng Yan, Le Zhuo, Zhen Li, Peng Gao, Zhanyu Ma, and Ming-Ming Cheng. 2025a · 2025
Closest in time.
IDEA-Bench: How Far are Generative Models from Professional Designing?. In Proceedings of the Computer Vision and Pattern Recognition Conference . 18541–18551
Chen Liang, Lianghua Huang, Jingwu Fang, Huanzhang Dou, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Junge Zhang, Xin Zhao, and Yu Liu. 2025 · 2025
Closest in time.
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
Yijing Lin, Mengqi Huang, Shuhan Zhuang, and Zhendong Mao. 2025 · 2025
Closest in time.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European Conference on Computer Vision . Springer, 38–55
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al · 2025
Closest in time.
Core: Context-regularized text embedding learning for text-to-image personalization. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 8377–8385
Feize Wu, Yun Pang, Junyi Zhang, Lianyu Pang, Jian Yin, Baoquan Zhao, Qing Li, and Xudong Mao. 2025 · 2025
Closest in time.
FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model
Jun Zhou, Jiahao Li, Zunnan Xu, Hanhui Li, Yiji Cheng, Fa-Ting Hong, Qin Lin, Qinglin Lu, and Xiaodan Liang. 2025 · 2025
Closest in time.