Fetching the paper…
Reading the bibliography…
We introduce ObjectAdd, a training-free diffusion modification method to add user-expected objects into user-specified area.
P. Soille et al. , Morphological image analysis: principles and applications . Springer, 1999, vol. 2
1999
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in Advances in Neural Information Processing Systems , 2017, pp. 6626–6637
2017
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in Neural Information Processing Systems , pp. 6840–6851, 2020
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning , 2021, pp. 8748–8763
2021
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-or, “Prompt-to-prompt image editing with cross-attention control,” in The Eleventh International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
W. Feng, X. He, T.-J. Fu, V. Jampani, A. R. Akula, P. Narayana, S. Basu, X. E. Wang, and W. Y. Wang, “Training-free structured diffusion guidance for compositional text-to-image synthesis,” in The Eleventh International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo et al. , “Improving image generation with better captions,” Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf , p. 3, 2023
2023
Earlier work this paper cites.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3836–3847
2023
Earlier work this paper cites.
Y. Kim, J. Lee, J.-H. Kim, J.-W. Ha, and J.-Y. Zhu, “Dense text-to-image generation with attention modulation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 7701–7711
2023
Cited alongside, same era.
O. Patashnik, D. Garibi, I. Azuri, H. Averbuch-Elor, and D. Cohen-Or, “Localizing object-level shape variations with text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023
2023
Cited alongside, same era.
H. Chefer, Y. Alaluf, Y. Vinker, L. Wolf, and D. Cohen-Or, “Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models,” ACM Transactions on Graphics , pp. 1–10, 2023
2023
Cited alongside, same era.
J. Xie, Y. Li, Y. Huang, H. Liu, W. Zhang, Y. Zheng, and M. Z. Shou, “Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 7452–7461
2023
Z. Pan, R. Gherardi, X. Xie, and S. Huang, “Effective real image editing with accelerated iterative diffusion inversion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023
2023
Later among the works it cites.
S. Zhengwentai, “clip-score: CLIP Score for PyTorch,” 2023, https://github.com/taited/clip-score
2023
Later among the works it cites.
2023
Later among the works it cites.
Midjourney, “Midjourney,” 2024, https://www.midjourney.com/home , Last accessed on 2024-2-27
2024
Closest in time.
Openai, “Chatgpt,” 2024, https://chat.openai.com/ , Last accessed on 2024-2-27
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2023
Cited alongside, same era.
C. Zhang, X. Chen, S. Chai, C. H. Wu, D. Lagun, T. Beeler, and F. De la Torre, “Iti-gen: Inclusive text-to-image generation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3969–3980
2023
Cited alongside, same era.
Y. Li, H. Liu, Q. Wu, F. Mu, J. Yang, J. Gao, C. Li, and Y. J. Lee, “Gligen: Open-set grounded text-to-image generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 22 511–22 521
2023
Cited alongside, same era.
T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 18 392–18 402
2023
Cited alongside, same era.
M. Cao, X. Wang, Z. Qi, Y. Shan, X. Qie, and Y. Zheng, “Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 22 560–22 570
2023
Cited alongside, same era.
G. Couairon, J. Verbeek, H. Schwenk, and M. Cord, “Diffedit: Diffusion-based semantic image editing with mask guidance,” in The Twelfth International Conference on Learning Representations , 2023
2023
Cited alongside, same era.
Closest in time.
Adobe, “Photoshop,” 2024, https://www.adobe.com/products/photoshop.html , Last accessed on 2024-2-27
2024
Closest in time.
M. Chen, I. Laina, and A. Vedaldi, “Training-free layout control with cross-attention guidance,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 5343–5353
2024
Closest in time.
Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma, “Investigating the catastrophic forgetting in multimodal large language model fine-tuning,” in Conference on Parsimony and Learning , 2024, pp. 202–227
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.