Fetching the paper…
Reading the bibliography…
This paper explores image modeling from the frequency space and introduces DCTdiff, an end-to-end diffusion generative paradigm that efficiently models images in the discrete cosine transform (DCT) space.
109. stochastic integral
Itô, K · 1944
Earlier work this paper cites.
Discrete cosine transform
Ahmed, N., Natarajan, T., and Rao, K. R · 1974
Earlier work this paper cites.
Reverse-time diffusion equation models
Anderson, B. D · 1982
Earlier work this paper cites.
The jpeg still picture compression standard
Wallace, G. K · 1991
Earlier work this paper cites.
Deflate compressed data format specification version 1.3
Deutsch, P · 1996
Earlier work this paper cites.
Origins of scaling in natural images
Ruderman, D. L · 1997
Earlier work this paper cites.
The multifractal structure of contrast changes in natural images: From sharp edges to textures
Turiel, A. and Parga, N · 2000
Earlier work this paper cites.
A fast scheme for image size change in the compressed domain
Dugad, R. and Ahuja, N · 2001
Earlier work this paper cites.
Digital image processing
Gonzalez, R. C · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Fast training of convolutional networks through ffts
Mathieu, M., Henaff, M., and LeCun, Y · 2014
Earlier work this paper cites.
Discrete cosine transform: algorithms, advantages, applications
Rao, K. R. and Yip, P · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Deep learning, 2016
Goodfellow, I · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Redmon, J · 2016
Earlier work this paper cites.
Improved techniques for training gans
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X · 2016
Earlier work this paper cites.
A downsampled variant of ImageNet as an alternative to the CIFAR datasets
Chrabaszcz, P., Loshchilov, I., and Hutter, F · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Fcnn: Fourier convolutional neural networks
Pratt, H., Williams, B., Coenen, F., and Zheng, Y · 2017
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
Learning long term dependencies via fourier recurrent units
Zhang, J., Lin, Y., Song, Z., and Dhillon, I · 2018
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Cited alongside, same era.
Improved precision and recall metric for assessing generative models
Kynkäänniemi, T., Karras, T., Laine, S., Lehtinen, J., and Aila, T · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Cited alongside, same era.
Mantra-net: Manipulation tracing network for detection and localization of image forgeries with anomalous features
Wu, Y., AbdAlmageed, W., and Natarajan, P · 2019
Cited alongside, same era.
Stargan v2: Diverse image synthesis for multiple domains
Choi, Y., Uh, Y., Yoo, J., and Ha, J.-W · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Understanding diffusion objectives as the elbo with simple data augmentation
Kingma, D. and Gao, R · 2023
Later among the works it cites.
Hierarchical vae with a diffusion-based vampprior
Kuzina, A. and Tomczak, J. M · 2023
Later among the works it cites.
Magic3d: High-resolution text-to-3d content creation
Lin, C.-H., Gao, J., Tang, L., Takikawa, T., Zeng, X., Huang, X., Kreis, K., Fidler, S., Liu, M.-Y., and Lin, T.-Y · 2023
Later among the works it cites.
Flow matching for generative modeling
Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M · 2023
Later among the works it cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Liu, X., Gong, C., and Liu, Q · 2023
Later among the works it cites.
Input perturbation reduces exposure bias in diffusion models
Ning, M., Sangineto, E., Porrello, A., Calderara, S., and Cucchiara, R · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language through a prism: A spectral approach for multiscale language representations
Tamkin, A., Jurafsky, D., and Goodman, N · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Generating images with sparse representations
Nash, C., Menick, J., Dieleman, S., and Battaglia, P · 2021
Cited alongside, same era.
Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models
Bao, F., Li, C., Zhu, J., and Zhang, B · 2022
Cited alongside, same era.
Fourier image transformer
Buchholz, T.-O. and Jug, F · 2022
Cited alongside, same era.
Later among the works it cites.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Later among the works it cites.
Wavelet diffusion models are fast and scalable image generators
Phung, H., Dao, Q., and Tran, A · 2023
Later among the works it cites.
Consistency models
Song, Y., Dhariwal, P., Chen, M., and Sutskever, I · 2023
Later among the works it cites.
Diffusion probabilistic model made slim
Yang, X., Zhou, D., Feng, J., and Wang, X · 2023
Later among the works it cites.
Diffusion is spectral autoregression, 2024
Dieleman, S · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Closest in time.
Esteves, C., Suhail, M., and Makadia, A · 2024
Closest in time.
Jpeg-lm: Llms as image generators with canonical codec representations
Han, X., Ghazvininejad, M., Koh, P. W., and Tsvetkov, Y · 2024
Closest in time.
Improved noise schedule for diffusion training
Hang, T. and Gu, S · 2024
Closest in time.
Rethinking fid: Towards a better evaluation metric for image generation
Jayasumana, S., Ramalingam, S., Veit, A., Glasner, D., Chakrabarti, A., and Kumar, S · 2024
Closest in time.
Analyzing and improving the training dynamics of diffusion models
Karras, T., Aittala, M., Lehtinen, J., Hellsten, J., Aila, T., and Laine, S · 2024
Closest in time.
Alleviating exposure bias in diffusion models through sampling with shifted time steps
Li, M., Qu, T., Yao, R., Sun, W., and Moens, M.-F · 2024
Closest in time.
Wavelets are all you need for autoregressive image generation
Mattar, W., Levy, I., Sharon, N., and Dekel, S · 2024
Closest in time.
Movie gen: A cast of media foundation models
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.-Y., Chuang, C.-Y., et al · 2024
Closest in time.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L · 2024
Closest in time.
Sa-solver: Stochastic adams solver for fast sampling of diffusion models
Xue, S., Yi, M., Luo, W., Zhang, S., Sun, J., Li, Z., and Ma, Z.-M · 2024
Closest in time.
Fast ode-based sampling for diffusion models in around 5 steps
Zhou, Z., Chen, D., Wang, C., and Chen, C · 2024
Closest in time.
Wavelet-based image tokenizer for vision transformers
Zhu, Z. and Soricut, R · 2024
Closest in time.
Nfig: Autoregressive image generation with next-frequency prediction
Huang, Z., Qiu, X., Ma, Y., Zhou, Y., Zhang, C., and Li, X · 2025
Closest in time.
Frequency autoregressive image generation with continuous tokens
Yu, H., Luo, H., Yuan, H., Rong, Y., and Zhao, F · 2025
Closest in time.