Fetching the paper…
Reading the bibliography…
With the advancement of image-to-image diffusion models guided by text, significant progress has been made in image editing.
Tweedie’s formula and selection bias
B. Efron · 2011
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
S. Barratt and R. Sharma · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Unbiased scene graph generation from biased training, 2020
K. Tang, Y. Niu, J. Huang, J. Shi, and H. Zhang · 2020
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Earlier work this paper cites.
Blended diffusion for text-driven editing of natural images
O. Avrahami, D. Lischinski, and O. Fried · 2022
Earlier work this paper cites.
ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers
Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, Q. Zhang, K. Kreis, M. Aittala, T. Aila, S. Laine, et al · 2022
Earlier work this paper cites.
Testing relational understanding in text-guided image generation
C. Conwell and T. Ullman · 2022
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion, 2022
R. Gal, Y. Alaluf, Y. Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or · 2022
Earlier work this paper cites.
Prompt-to-prompt image editing with cross attention control
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or · 2022
Earlier work this paper cites.
J. Li, D. Li, C. Xiong, and S. C. H. Hoi · 2022
Earlier work this paper cites.
Sdedit: Guided image synthesis and editing with stochastic differential equations, 2022
C. Meng, Y. He, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
Dalle-2 is seeing double: flaws in word-to-concept mapping in text2image models
R. Rassin, S. Ravfogel, and Y. Goldberg · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Cited alongside, same era.
LAION-5B: an open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev · 2022
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Later among the works it cites.
Localizing object-level shape variations with text-to-image diffusion models, 2023
O. Patashnik, D. Garibi, I. Azuri, H. Averbuch-Elor, and D. Cohen-Or · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation, 2023
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman · 2023
Later among the works it cites.
How to bridge the gap between modalities: A comprehensive survey on multimodal large language model, 2023
S. Song, X. Li, S. Li, S. Zhao, J. Yu, J. Ma, X. Mao, and W. Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, et al · 2022
Cited alongside, same era.
Scene graph generation: A comprehensive survey
G. Zhu, L. Zhang, Y. Jiang, Y. Dang, H. Hou, P. Shen, M. Feng, X. Zhao, Q. Miao, S. A. A. Shah, et al · 2022
Cited alongside, same era.
Openflamingo: An open-source framework for training large autoregressive vision-language models
A. Awadalla, I. Gao, J. Gardner, J. Hessel, Y. Hanafy, W. Zhu, K. Marathe, Y. Bitton, S. Gadre, S. Sagawa, et al · 2023
Cited alongside, same era.
Instructpix2pix: Learning to follow image editing instructions, 2023
T. Brooks, A. Holynski, and A. A. Efros · 2023
Cited alongside, same era.
Erasing concepts from diffusion models
R. Gandikota, J. Materzynska, J. Fiotto-Kaufman, and D. Bau · 2023
Cited alongside, same era.
simple diffusion: End-to-end diffusion for high resolution images
E. Hoogeboom, J. Heek, and T. Salimans · 2023
Cited alongside, same era.
Improving faithfulness for vision transformers
L. Hu, Y. Liu, N. Liu, M. Huai, L. Sun, and D. Wang · 2023
Cited alongside, same era.
Dreaminpainter: Text-guided subject-driven image inpainting with diffusion models, 2023
S. Xie, Y. Zhao, Z. Xiao, K. C. K. Chan, Y. Li, Y. Xu, K. Zhang, and T. Hou · 2023
Later among the works it cites.
Inpaint anything: Segment anything meets image inpainting
T. Yu, R. Feng, R. Feng, J. Liu, X. Jin, W. Zeng, and Z. Chen · 2023
Later among the works it cites.
Erasing concepts from text-to-image diffusion models with few-shot unlearning
M. Fuchi and T. Takagi · 2024
Closest in time.
Unified concept editing in diffusion models
R. Gandikota, H. Orgad, Y. Belinkov, J. Materzyńska, and D. Bau · 2024
Closest in time.
Anomalydiffusion: Few-shot anomaly image generation with diffusion model
T. Hu, J. Zhang, R. Yi, Y. Du, X. Chen, L. Liu, Y. Wang, and C. Wang · 2024
Closest in time.
Diffusion model-based image editing: A survey, 2024
Y. Huang, J. Huang, Y. Liu, M. Yan, J. Lv, J. Liu, W. Xiong, H. Zhang, S. Chen, and L. Cao · 2024
Closest in time.
Towards understanding cross and self-attention in stable diffusion for text-guided image editing, 2024
B. Liu, C. Wang, T. Cao, K. Jia, and J. Huang · 2024
Closest in time.
Shape-guided diffusion with inside-outside attention
D. H. Park, G. Luo, C. Toste, S. Azadi, X. Liu, M. Karalashvili, A. Rohrbach, and T. Darrell · 2024
Closest in time.
Magicbrush: A manually annotated dataset for instruction-guided image editing, 2024
K. Zhang, L. Mo, W. Chen, H. Sun, and Y. Su · 2024
Closest in time.