Fetching the paper…
Reading the bibliography…
While likelihood-based generative models, particularly diffusion and autoregressive models, have achieved remarkable fidelity in visual generation, the maximum likelihood estimation (MLE) objective, which minimizes the forward KL divergence, inherently suffers from a mode-covering tendency that limits the generation quality under limited model capacity.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Pattern recognition and machine learning , volume 4
Bishop, C. M. and Nasrabadi, N. M · 2006
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Density estimation using real nvp
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2016
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
f-gan: Training generative neural samplers using variational divergence minimization
Nowozin, S., Cseke, B., and Tomioka, R · 2016
Earlier work this paper cites.
Pixel recurrent neural networks
Van Den Oord, A., Kalchbrenner, N., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Barratt, S. and Sharma, R · 2018
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
Brock, A · 2018
Earlier work this paper cites.
Neural ordinary differential equations
Chen, R. T., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Implicit generation and modeling with energy based models
Du, Y. and Mordatch, I · 2019
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Ghazvininejad, M., Levy, O., Liu, Y., and Zettlemoyer, L · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Training generative adversarial networks with limited data
Karras, T., Aittala, M., Hellsten, J., Laine, S., Lehtinen, J., and Aila, T · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and Van Den Berg, R · 2021
Earlier work this paper cites.
Diffusion models beat GANs on image synthesis
Dhariwal, P. and Nichol, A. Q · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2021
Cited alongside, same era.
Variational diffusion models
Kingma, D. P., Salimans, T., Poole, B., and Ho, J · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
Score-based generative modeling in latent space
Vahdat, A., Kreis, K., and Kautz, J · 2021
Cited alongside, same era.
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2021
Cited alongside, same era.
Language model beats diffusion–tokenizer is key to visual generation
Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Birodkar, V., Gupta, A., Gu, X., et al · 2023
Later among the works it cites.
Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models
Bao, F., Xiang, C., Yue, G., He, G., Zhu, H., Zheng, K., Zhao, M., Liu, S., Wang, Y., and Zhu, J · 2024
Later among the works it cites.
Video generation models as world simulators. 2024
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., et al · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Later among the works it cites.
Pagoda: Progressive growing of a one-step generator from a low-resolution diffusion teacher
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers
Balaji, Y., Nah, S., Huang, X., Vahdat, A., Song, J., Zhang, Q., Kreis, K., Aittala, M., Aila, T., Laine, S., et al · 2022
Cited alongside, same era.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al · 2022
Cited alongside, same era.
Scalable adaptive computation for iterative generation
Jabri, A., Fleet, D., and Chen, T · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Autoregressive image generation using residual quantization
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W.-S · 2022
Cited alongside, same era.
Kim, D., Lai, C.-H., Liao, W.-H., Takida, Y., Murata, N., Uesaka, T., Mitsufuji, Y., and Ermon, S · 2024
Later among the works it cites.
Understanding diffusion objectives as the elbo with simple data augmentation
Kingma, D. and Gao, R · 2024
Later among the works it cites.
Autoregressive image generation without vector quantization
Li, T., Tian, Y., Li, H., Deng, M., and He, K · 2024
Later among the works it cites.
Open-magvit2: An open-source project toward democratizing auto-regressive visual generation
Luo, Z., Shi, F., Ge, Y., Yang, Y., Wang, L., and Shan, Y · 2024
Later among the works it cites.
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers
Ma, N., Goldstein, M., Albergo, M. S., Boffi, N. M., Vanden-Eijnden, E., and Xie, S · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Simple and effective masked diffusion language models
Sahoo, S. S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J. T., Rush, A., and Kuleshov, V · 2024
Later among the works it cites.
Simplified and generalized masked diffusion for discrete data
Shi, J., Han, K., Wang, Z., Doucet, A., and Titsias, M. K · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., and Yuan, Z · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L · 2024
Later among the works it cites.
Givt: Generative infinite-vocabulary transformers
Tschannen, M., Eastwood, C., and Mentzer, F · 2024
Later among the works it cites.
Diffusion model alignment using direct preference optimization
Wallace, B., Dang, M., Rafailov, R., Zhou, L., Lou, A., Purushwalkam, S., Ermon, S., Xiong, C., Joty, S., and Naik, N · 2024
Later among the works it cites.
Maskbit: Embedding-free image generation via bit tokens
Weber, M., Yu, L., Yu, Q., Deng, X., Shen, X., Cremers, D., and Chen, L.-C · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z · 2024
Later among the works it cites.
Disco-diff: Enhancing continuous diffusion models with discrete latents
Xu, Y., Corso, G., Jaakkola, T., Vahdat, A., and Kreis, K · 2024
Later among the works it cites.
Improved distribution matching distillation for fast image synthesis
Yin, T., Gharbi, M., Park, T., Zhang, R., Shechtman, E., Durand, F., and Freeman, W. T · 2024
Later among the works it cites.
Zheng, K., Chen, Y., Mao, H., Liu, M.-Y., Zhu, J., and Zhang, Q · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
The gan is dead; long live the gan! a modern gan baseline
Huang, Y., Gokaslan, A., Kuleshov, V., and Tompkin, J · 2025
Closest in time.
Beyond next-token: Next-x prediction for autoregressive visual generation
Ren, S., Yu, Q., He, J., Shen, X., Yuille, A., and Chen, L.-C · 2025
Closest in time.
The diffusion duality
Sahoo, S. S., Deschenaux, J., Gokaslan, A., Wang, G., Chiu, J. T., and Kuleshov, V · 2025
Closest in time.
Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models
Yao, J., Yang, B., and Wang, X · 2025
Closest in time.