Fetching the paper…
Reading the bibliography…
Diffusion models have demonstrated remarkable performance in generating unimodal data across various tasks, including image, video, and text generation.
Reverse-time diffusion equation models
Anderson, B. D · 1982
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A. and Dayan, P · 2005
Earlier work this paper cites.
Reversibility and stochastic networks
Kelly, F. P · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Interpretation and generalization of score matching
Lyu, S · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
A continuous time framework for discrete denoising models
Campbell, A., Benton, J., De Bortoli, V., Rainforth, T., Deligiannidis, G., and Doucet, A · 2022
Earlier work this paper cites.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Earlier work this paper cites.
Riemannian score-based generative modelling
De Bortoli, V., Mathieu, E., Hutchinson, M., Thornton, J., Teh, Y. W., and Doucet, A · 2022
Earlier work this paper cites.
Video diffusion models
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Earlier work this paper cites.
Unified discrete diffusion for simultaneous vision-language generation
Hu, M., Zheng, C., Zheng, H., Cham, T.-J., Wang, C., Yang, Z., Tao, D., and Suganthan, P. N · 2022
Earlier work this paper cites.
Riemannian diffusion models
Huang, C.-W., Aghajohari, M., Bose, J., Panangaden, P., and Courville, A. C · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Earlier work this paper cites.
Stasy: Score-based tabular data synthesis
Kim, J., Lee, C., and Park, N · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Cited alongside, same era.
Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wu, Y., Wang, Z., Kwok, J., Luo, P., Lu, H., et al · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S · 2023
Cited alongside, same era.
Segment anything
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Cited alongside, same era.
Applying guidance in a limited interval improves sample and distribution quality in diffusion models
Kynkäänniemi, T., Aittala, M., Karras, T., Laine, S., Aila, T., and Lehtinen, J · 2024
Later among the works it cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2024
Later among the works it cites.
Discrete diffusion modeling by estimating the ratios of the data distribution
Lou, A., Meng, C., and Ermon, S · 2024
Later among the works it cites.
Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., and Kembhavi, A · 2024
Later among the works it cites.
scdiffusion: conditional generation of high-quality single-cell data using diffusion model
Luo, E., Hao, M., Wei, L., and Zhang, X · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tabddpm: Modelling tabular data with diffusion models
Kotelnikov, A., Baranchuk, D., Rubachev, I., and Babenko, A · 2023
Cited alongside, same era.
Codi: Co-evolving contrastive diffusion models for mixed-type tabular synthesis
Lee, C., Kim, J., and Park, N · 2023
Cited alongside, same era.
Goggle: Generative modelling for tabular data by learning relational structure
Liu, T., Qian, Z., Berrevoets, J., and van der Schaar, M · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation
Ruan, L., Ma, Y., Yang, H., He, H., Liu, B., Fu, J., Yuan, N. J., Jin, Q., and Guo, B · 2023
Cited alongside, same era.
De novo design of protein structure and function with rfdiffusion
Watson, J. L., Juergens, D., Bennett, N. R., Trippe, B. L., Yim, J., Eisenach, H. E., Ahern, W., Borst, A. J., Ragotte, R. J., Milles, L. F., et al · 2023
Cited alongside, same era.
Versatile diffusion: Text, images and variations all in one diffusion model
Xu, X., Wang, Z., Zhang, G., Wang, K., and Shi, H · 2023
Cited alongside, same era.
Meta, C · 2024
Later among the works it cites.
Unlocking guidance for discrete state-space diffusion and flow models
Nisonoff, H., Xiong, J., Allenspach, S., and Listgarten, J · 2024
Later among the works it cites.
Ren, Y., Chen, H., Rotskoff, G. M., and Ying, L · 2024
Later among the works it cites.
Simple and effective masked diffusion language models
Sahoo, S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J., Rush, A., and Kuleshov, V · 2024
Later among the works it cites.
Simple guidance mechanisms for discrete diffusion models
Schiff, Y., Sahoo, S. S., Phung, H., Wang, G., Boshar, S., Dalla-torre, H., de Almeida, B. P., Rush, A., Pierrot, T., and Kuleshov, V · 2024
Later among the works it cites.
Simplified and generalized masked diffusion for discrete data
Shi, J., Han, K., Wang, Z., Doucet, A., and Titsias, M · 2024
Later among the works it cites.
Flowllm: Flow matching for material generation with large language models as base distributions
Sriram, A., Miller, B. K., Chen, R. T., and Wood, B. M · 2024
Later among the works it cites.
Jetformer: An autoregressive generative model of raw images and text
Tschannen, M., Pinto, A. S., and Kolesnikov, A · 2024
Later among the works it cites.
Janus: Decoupling visual encoding for unified multimodal understanding and generation
Wu, C., Chen, X., Wu, Z., Ma, Y., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., Ruan, C., et al · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z · 2024
Later among the works it cites.
Discrete-state continuous-time diffusion for graph generation
Xu, Z., Qiu, R., Chen, Y., Chen, H., Fan, X., Pan, M., Zeng, Z., Das, M., and Tong, H · 2024
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Zhou, C., Yu, L., Babu, A., Tirumala, K., Yasunaga, M., Shamis, L., Kahn, J., Ma, X., Zettlemoyer, L., and Levy, O · 2024
Later among the works it cites.
Stiefel flow matching for moment-constrained structure elucidation
Cheng, A., Lo, A., Lee, K. L. K., Miret, S., and Aspuru-Guzik, A · 2025
Closest in time.
Simulating 500 million years of evolution with a language model
Hayes, T., Rao, R., Akin, H., Sofroniew, N. J., Oktay, D., Lin, Z., Verkuil, R., Tran, V. Q., Deaton, J., Wiggert, M., et al · 2025
Closest in time.
Pyramidal flow matching for efficient video generative modeling
Jin, Y., Sun, Z., Li, N., Xu, K., Jiang, H., Zhuang, N., Huang, Q., Song, Y., Mu, Y., and Lin, Z · 2025
Closest in time.
Layerdag: A layerwise autoregressive diffusion model for directed acyclic graph generation
Li, M., Shitole, V., Chien, E., Man, C., Wang, Z., Sridharan, S., Zhang, Y., Krishna, T., and Li, P · 2025
Closest in time.
Your absorbing discrete diffusion secretly models the conditional distributions of clean data
Ou, J., Nie, S., Xue, K., Zhu, F., Sun, J., Li, Z., and Li, C · 2025
Closest in time.
Variational schrödinger momentum diffusion
Rojas, K., Tian, Y., Tao, M., Nevmyvaka, Y., and Deng, W · 2025
Closest in time.
Unified multimodal discrete diffusion
Swerdlow, A., Prabhudesai, M., Gandhi, S., Pathak, D., and Fragkiadaki, K · 2025
Closest in time.