Fetching the paper…
Reading the bibliography…
The expressive power and computational complexity of deep visual generative models, such as flow-based and autoregressive (AR) models, have gained considerable interest for their wide-ranging applications in generative tasks.
Time, hardware, and uniformity
Barrington, D. M. and Immerman, N. (1994) · 1994
Earlier work this paper cites.
Computational complexity: a modern approach
Arora, S. and Barak, B. (2009) · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A. (2020) · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. (2020) · 2011
Earlier work this paper cites.
Tutorial on variational autoencoders
Doersch, C. (2016) · 2016
Earlier work this paper cites.
Glow: Generative flow with invertible 1x1 convolutions
Kingma, D. P. and Dhariwal, P. (2018) · 2018
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2020) · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P. (2020) · 2020
Earlier work this paper cites.
Sit: Self-supervised vision transformer
Atito, S., Awais, M., and Kittler, J. (2021) · 2021
Earlier work this paper cites.
Sublinear time algorithm for online weighted bipartite matching
Hu, H., Song, Z., Tao, R., Xu, Z., Yin, J., and Zhuo, D. (2022) · 2022
Earlier work this paper cites.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C. (2022) · 2022
Cited alongside, same era.
Saturated transformers are constant-depth threshold circuits
Merrill, W., Sabharwal, A., and Smith, N. A. (2022) · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022) · 2022
Cited alongside, same era.
Fast attention requires bounded entries
Alman, J. and Song, Z. (2023) · 2023
Cited alongside, same era.
All are worth words: A vit backbone for diffusion models
Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., and Zhu, J. (2023) · 2023
Cited alongside, same era.
simple diffusion: End-to-end diffusion for high resolution images
Chiang, D. (2024) · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al. (2024) · 2024
Later among the works it cites.
Pyramidal flow matching for efficient video generative modeling
Jin, Y., Sun, Z., Li, N., Xu, K., Jiang, H., Zhuang, N., Huang, Q., Song, Y., Mu, Y., and Lin, Z. (2024) · 2024
Later among the works it cites.
A logic for expressing log-precision transformers
Merrill, W. and Sabharwal, A. (2024) · 2024
Later among the works it cites.
Flowar: Scale-wise autoregressive image generation meets flow matching
Ren, S., Yu, Q., He, J., Shen, X., Yuille, A., and Chen, L.-C. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hoogeboom, E., Heek, J., and Salimans, T. (2023) · 2023
Cited alongside, same era.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Liu, X., Gong, C., and Liu, Q. (2023) · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S. (2023) · 2023
Cited alongside, same era.
Dolfin: Diffusion layout transformers without autoencoder
Wang, Y., Chen, Z., Zhong, L., Ding, Z., Sha, Z., and Tu, Z. (2023) · 2023
Cited alongside, same era.
Fast rope attention: Combining the polynomial method and fast fourier transform
Alman, J. and Song, Z. (2024a)
Cited in the paper.
The fine-grained complexity of gradient computation for training large language models
Alman, J. and Song, Z. (2024b)
Cited in the paper.
How to capture higher-order correlations? generalizing matrix softmax attention to kronecker computation
Alman, J. and Song, Z. (2024c)
Cited in the paper.
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L. (2024) · 2024
Later among the works it cites.
When can we solve the weighted low rank approximation problem in truly subquadratic time?
Li, C., Liang, Y., Shi, Z., and Song, Z. (2025) · 2025
Closest in time.
Looped relu mlps may be all you need as practical programmable computers
Liang, Y., Sha, Z., Shi, Z., Song, Z., and Zhou, Y. (2025) · 2025
Closest in time.
Fast and efficient matching algorithm with deadline instances
Song, Z., Wang, W., Yin, C., and Yin, J. (2025) · 2025
Closest in time.