Fetching the paper…
Reading the bibliography…
Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis.
Generating diverse high-fidelity images with vq-vae-2, 2019b
Razavi, A., van den Oord, A., and Vinyals, O · 1906
Earlier work this paper cites.
Histograms of oriented gradients for human detection
Dalal, N. and Triggs, B · 2005
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. R · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2010
Earlier work this paper cites.
Denoising diffusion implicit models, 2022
Song, J., Meng, C., and Ermon, S · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Generating images with perceptual similarity metrics based on deep networks
Dosovitskiy, A. and Brox, T · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Johnson, J., Alahi, A., and Fei-Fei, L · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
Larsen, A. B. L., Sønderby, S. K., Larochelle, H., and Winther, O · 2016
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Van den Oord, A., Kalchbrenner, N., Espeholt, L., Vinyals, O., Graves, A., et al · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C. P., Glorot, X., Botvinick, M. M., Mohamed, S., and Lerchner, A · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks, 2018
Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
McInnes, L., Healy, J., and Melville, J · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric, 2018
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Earlier work this paper cites.
Improved precision and recall metric for assessing generative models
Kynkäänniemi, T., Karras, T., Laine, S., Lehtinen, J., and Aila, T · 2019
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models, 2021
Nichol, A. and Dhariwal, P · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Regularizing generative adversarial networks under limited data, 2021
Tseng, H.-Y., Jiang, L., Liu, C., Yang, M.-H., and Yang, W · 2021
Cited alongside, same era.
Score-based generative modeling in latent space, 2021
Vahdat, A., Kreis, K., and Kautz, J · 2021
Cited alongside, same era.
Learning mixtures of gaussians using the ddpm objective
Shah, K., Chen, S., and Klivans, A · 2023
Later among the works it cites.
Attention is all you need, 2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Later among the works it cites.
Nearly d-linear convergence bounds for diffusion models via stochastic localization
Benton, J., Bortoli, V., Doucet, A., and Deligiannidis, G · 2024
Later among the works it cites.
Cfg++: Manifold-constrained classifier free guidance for diffusion models
Chung, H., Kim, J., Park, G. Y., Nam, H., and Ye, J. C · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2021
Cited alongside, same era.
Maskgit: Masked generative image transformer, 2022
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Cited alongside, same era.
Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions
Chen, S., Chewi, S., Li, J., Li, Y., Salim, A., and Zhang, A. R · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Cited alongside, same era.
Autoregressive image generation using residual quantization
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W.-S · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Deng, C., Zh, D., Li, K., Guan, S., and Fan, H · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Later among the works it cites.
Learning mixtures of gaussians using diffusion models
Gatmiry, K., Kelner, J., and Lee, H · 2024
Later among the works it cites.
Geneval: An object-focused framework for evaluating text-to-image alignment
Ghosh, D., Hajishirzi, H., and Schmidt, L · 2024
Later among the works it cites.
Classification done right for vision-language pre-training
Huang, Z., Ye, Q., Kang, B., Feng, J., and Fan, H · 2024
Later among the works it cites.
Guiding a diffusion model with a bad version of itself
Karras, T., Aittala, M., Kynkäänniemi, T., Lehtinen, J., Aila, T., and Laine, S · 2024
Later among the works it cites.
Applying guidance in a limited interval improves sample and distribution quality in diffusion models
Kynkäänniemi, T., Aittala, M., Karras, T., Laine, S., Aila, T., and Lehtinen, J · 2024
Later among the works it cites.
Customize your visual autoregressive recipe with set autoregressive modeling
Liu, W., Zhuo, L., Xin, Y., Xia, S., Gao, P., and Yue, X · 2024
Later among the works it cites.
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers
Ma, N., Goldstein, M., Albergo, M. S., Boffi, N. M., Vanden-Eijnden, E., and Xie, S · 2024
Later among the works it cites.
Tokenflow: Unified image tokenizer for multimodal understanding and generation
Qu, L., Zhang, H., Liu, Y., Wang, X., Jiang, Y., Gao, Y., Ye, H., Du, D. K., Yuan, Z., and Wu, X · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., and Yuan, Z · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction, 2024
Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L · 2024
Later among the works it cites.
Maskbit: Embedding-free image generation via bit tokens
Weber, M., Yu, L., Yu, Q., Deng, X., Shen, X., Cremers, D., and Chen, L.-C · 2024
Later among the works it cites.
Vila-u: a unified foundation model integrating visual understanding and generation
Wu, Y., Zhang, Z., Chen, J., Tang, H., Li, D., Fang, Y., Zhu, L., Xie, E., Yin, H., Yi, L., et al · 2024
Later among the works it cites.
Language-guided image tokenization for generation
Zha, K., Yu, L., Fathi, A., Ross, D. A., Schmid, C., Katabi, D., and Gu, X · 2024
Later among the works it cites.
Scaling the codebook size of vqgan to 100,000 with a utilization rate of 99%
Zhu, L., Wei, F., Lu, Y., and Chen, D · 2024
Later among the works it cites.
Visual generation without guidance
Chen, H., Jiang, K., Zheng, K., Chen, J., Su, H., and Zhu, J · 2025
Closest in time.
Robust latent matters: Boosting image generation with sampling error synthesis
Qiu, K., Li, X., Kuen, J., Chen, H., Xu, X., Gu, J., Luo, Y., Raj, B., Lin, Z., and Savvides, M · 2025
Closest in time.
Givt: Generative infinite-vocabulary transformers
Tschannen, M., Eastwood, C., and Mentzer, F · 2025
Closest in time.
Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models
Yao, J. and Wang, X · 2025
Closest in time.
Studying classifier (-free) guidance from a classifier-centric perspective
Zhao, X. and Schwing, A. G · 2025
Closest in time.