Fetching the paper…
Reading the bibliography…
Despite the remarkable performance of text-to-image diffusion models in image generation tasks, recent studies have raised the issue that generated images sometimes cannot capture the intended semantic contents of the text prompts, which phenomenon is often called semantic misalignment.
The concave-convex procedure (cccp)
A. L. Yuille and A. Rangarajan · 2001
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
On a model of associative memory with huge storage capacity
M. Demircigil, J. Heusel, M. Löwe, S. Upgang, and F. Vermet · 2017
Earlier work this paper cites.
Introvae: Introspective variational autoencoders for photographic image synthesis
H. Huang, R. He, Z. Sun, T. Tan, et al · 2018
Earlier work this paper cites.
Dense associative memory is robust to adversarial inputs
D. Krotov and J. Hopfield · 2018
Earlier work this paper cites.
Implicit generation and generalization in energy-based models
Y. Du and I. Mordatch · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
T. Karras, S. Laine, and T. Aila · 2019
Earlier work this paper cites.
Learning non-convergent non-persistent short-run mcmc toward energy-based model
E. Nijkamp, M. Hill, S.-C. Zhu, and Y. N. Wu · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Y. Song and S. Ermon · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Stargan v2: Diverse image synthesis for multiple domains
Y. Choi, Y. Uh, J. Yoo, and J.-W. Ha · 2020
Earlier work this paper cites.
Compositional visual generation with energy based models
Y. Du, S. Li, and I. Mordatch · 2020
Earlier work this paper cites.
Improved contrastive divergence training of energy based models
Y. Du, S. Li, J. Tenenbaum, and I. Mordatch · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Hopfield networks is all you need
H. Ramsauer, B. Schäfl, J. Lehner, P. Seidl, M. Widrich, T. Adler, L. Gruber, M. Holzleitner, M. Pavlović, G. K. Sandve, et al · 2020
Cited alongside, same era.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Cited alongside, same era.
Repaint: Inpainting using denoising diffusion probabilistic models
A. Lugmayr, M. Danelljan, A. Romero, F. Yu, R. Timofte, and L. Van Gool · 2022
Later among the works it cites.
Null-text inversion for editing real images using guided diffusion models
R. Mokady, A. Hertz, K. Aberman, Y. Pritch, and D. Cohen-Or · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. P. Kingma, T. Salimans, B. Poole, and J. Ho · 2021
Cited alongside, same era.
Sdedit: Image synthesis and editing with stochastic differential equations
C. Meng, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon · 2021
Cited alongside, same era.
Controllable and compositional generation with latent-space energy-based models
W. Nie, A. Vahdat, and A. Anandkumar · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Training-free structured diffusion guidance for compositional text-to-image synthesis
W. Feng, X. He, T.-J. Fu, V. Jampani, A. Akula, P. Narayana, S. Basu, X. E. Wang, and W. Y. Wang · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Cited alongside, same era.
Later among the works it cites.
Splicing vit features for semantic appearance transfer
N. Tumanyan, O. Bar-Tal, S. Bagon, and T. Dekel · 2022
Later among the works it cites.
Plug-and-play diffusion features for text-driven image-to-image translation
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel · 2022
Later among the works it cites.
Unifying diffusion models’ latent space, with applications to cyclediffusion and guidance
C. H. Wu and F. De la Torre · 2022
Later among the works it cites.
Smartbrush: Text and shape guided object inpainting with diffusion model
S. Xie, Z. Zhang, Z. Lin, T. Hinz, and K. Zhang · 2022
Later among the works it cites.
Geometry of Deep Learning
J. C. Ye · 2022
Later among the works it cites.
Sega: Instructing diffusion using semantic dimensions
M. Brack, F. Friedrich, D. Hintersdorf, L. Struppek, P. Schramowski, and K. Kersting · 2023
Closest in time.
Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models
H. Chefer, Y. Alaluf, Y. Vinker, L. Wolf, and D. Cohen-Or · 2023
Closest in time.
B. Hoover, Y. Liang, B. Pham, R. Panda, H. Strobelt, D. H. Chau, M. J. Zaki, and D. Krotov · 2023
Closest in time.
Zero-shot image-to-image translation
G. Parmar, K. K. Singh, R. Zhang, Y. Li, J. Lu, and J.-Y. Zhu · 2023
Closest in time.