Fetching the paper…
Reading the bibliography…
Recently, the strong latent Diffusion Probabilistic Model (DPM) has been applied to high-quality Text-to-Image (T2I) generation (e.g., Stable Diffusion), by injecting the encoded target text prompt into the gradually denoised diffusion image generator.
Digital image processing
R. C. Gonzales and P. Wintz · 1987
Earlier work this paper cites.
The nature of statistical learning theory
V. Vapnik · 1999
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Metrics for evaluating 3d medical image segmentation: analysis, selection, and tool
A. A. Taha and A. Hanbury · 2015
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
J. Hessel, A. Holtzman, M. Forbes, R. Le Bras, and Y. Choi · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2021
Earlier work this paper cites.
Generate your counterfactuals: Towards controlled counterfactual generation for text
N. Madaan, I. Padhi, N. Panwar, and D. Saha · 2021
Earlier work this paper cites.
Oodgan: Generative adversarial network for out-of-domain data generation
P. Marek, V. I. Naik, A. Goyal, and V. Auvray · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Do long-range language models actually use long-range context?
S. Sun, K. Krishna, A. Mattarella-Micke, and M. Iyyer · 2021
Earlier work this paper cites.
Counterfactual memorization in neural language models
C. Zhang, D. Ippolito, K. Lee, M. Jagielski, F. Tramèr, and N. Carlini · 2021
Earlier work this paper cites.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Y. Balaji, S. Nah, X. Huang, A. Vahdat, J. Song, K. Kreis, M. Aittala, T. Aila, S. Laine, B. Catanzaro, et al · 2022
Cited alongside, same era.
Rankgen: Improving text generation with large ranking models
K. Krishna, Y. Chang, J. Wieting, and M. Iyyer · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Cited alongside, same era.
Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps
C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu · 2022
Cited alongside, same era.
Image segmentation using text and image prompts
T. Lüddecke and A. Ecker · 2022
Cited alongside, same era.
Scaling up gans for text-to-image synthesis
M. Kang, J.-Y. Zhu, R. Zhang, J. Park, E. Shechtman, S. Paris, and T. Park · 2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Later among the works it cites.
Magic3d: High-resolution text-to-3d content creation
C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y. Liu, and T.-Y. Lin · 2023
Later among the works it cites.
Lost in the middle: How language models use long contexts
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang · 2023
Later among the works it cites.
Dreamix: Video diffusion models are general video editors
E. Molad, E. Horwitz, D. Valevski, A. R. Acha, Y. Matias, Y. Pritch, Y. Leviathan, and Y. Hoshen · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Cited alongside, same era.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, et al · 2022
Cited alongside, same era.
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman · 2023
Later among the works it cites.
Freeu: Free lunch in diffusion u-net
C. Si, Z. Huang, Y. Jiang, and Z. Liu · 2023
Later among the works it cites.
What the daam: Interpreting stable diffusion using cross attention
R. Tang, A. Pandey, Z. Jiang, G. Yang, K. Kumar, J. Lin, and F. Ture · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis · 2023
Later among the works it cites.
Diffusion probabilistic model made slim
X. Yang, D. Zhou, J. Feng, and X. Wang · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis
J. Chen, Y. Jincheng, G. Chongjian, L. Yao, E. Xie, Z. Wang, J. Kwok, P. Luo, H. Lu, and Z. Li · 2024
Closest in time.
Anydoor: Zero-shot object-level image customization
Y. L. Y. S. D. Z. H. Z. Xi Chen, Lianghua Huang · 2024
Closest in time.
Cross-attention makes inference cumbersome in text-to-image diffusion models
W. Zhang, H. Liu, J. Xie, F. Faccio, M. Z. Shou, and J. Schmidhuber · 2024
Closest in time.
Photomaker: Customizing realistic human photos via stacked id embedding
X. W. Z. Q. M.-M. C. Y. S. Zhen Li, Mingdeng Cao · 2024
Closest in time.