Fetching the paper…
Reading the bibliography…
We introduce a novel state-space architecture for diffusion models, effectively harnessing spatial and frequency information to enhance the inductive bias towards local features in input images for image generation tasks.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop
F. Yu, Y. Zhang, S. Song, A. Seff, and J. Xiao · 2015
Earlier work this paper cites.
Introvae: Introspective variational autoencoders for photographic image synthesis
H. Huang, z. li, R. He, Z. Sun, and T. Tan · 2018
Earlier work this paper cites.
Large scale GAN training for high fidelity natural image synthesis
A. Brock, J. Donahue, and K. Simonyan · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
T. Karras, S. Laine, and T. Aila · 2019
Earlier work this paper cites.
Improved precision and recall metric for assessing generative models
T. Kynkäänniemi, T. Karras, S. Laine, J. Lehtinen, and T. Aila · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Y. Song and S. Ermon · 2019
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Training generative adversarial networks with limited data
T. Karras, M. Aittala, J. Hellsten, S. Laine, J. Lehtinen, and T. Aila · 2020
Earlier work this paper cites.
Reliable fidelity and diversity metrics for generative models
M. F. Naeem, S. J. Oh, Y. Uh, Y. Choi, and J. Yoo · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Earlier work this paper cites.
A fast jpeg image compression algorithm based on dct
W. Xiao, N. Wan, A. Hong, and X. Chen · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2021
Earlier work this paper cites.
Score-based generative modeling in latent space
A. Vahdat, K. Kreis, and J. Kautz · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Earlier work this paper cites.
On the parameterization and initialization of diagonal state space models
A. Gu, K. Goel, A. Gupta, and C. Ré · 2022
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
A. Gu, K. Goel, and C. Ré · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, et al · 2022
Cited alongside, same era.
Nommer: Nominate synergistic context in vision transformer for visual recognition
H. Liu, X. Jiang, X. Li, Z. Bao, D. Jiang, and B. Ren · 2022
Cited alongside, same era.
Scalable diffusion models with transformers. 2023 ieee
W. S. Peebles and S. Xie · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Stylegan-xl: Scaling stylegan to large diverse datasets
A. Sauer, K. Schwarz, and A. Geiger · 2022
Cited alongside, same era.
Tackling the generative learning trilemma with denoising diffusion gans
Zamba: A compact 7b ssm hybrid model
P. Glorioso, Q. Anthony, Y. Tokpanov, J. Whittington, J. Pilault, A. Ibrahim, and B. Millidge · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
A. Gu and T. Dao · 2024
Closest in time.
Mambair: A simple baseline for image restoration with state-space model
H. Guo, J. Li, T. Dai, Z. Ouyang, X. Ren, and S.-T. Xia · 2024
Closest in time.
Zigma: A dit-style zigzag mamba diffusion model
V. T. Hu, S. A. Baumann, M. Gui, O. Grebenkova, P. Ma, J. Fischer, and B. Ommer · 2024
Closest in time.
Localmamba: Visual state space model with windowed selective scan
T. Huang, X. Pei, S. You, F. Wang, C. Qian, and C. Xu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Xiao, K. Kreis, and A. Vahdat · 2022
Cited alongside, same era.
Wavegan: Frequency-aware gan for high-fidelity few-shot image generation
M. Yang, Z. Wang, Z. Chi, and W. Feng · 2022
Cited alongside, same era.
Wave-vit: Unifying wavelet and transformers for visual representation learning
T. Yao, Y. Pan, Y. Li, C.-W. Ngo, and T. Mei · 2022
Cited alongside, same era.
All are worth words: A vit backbone for diffusion models
F. Bao, S. Nie, K. Xue, Y. Cao, C. Li, H. Su, and J. Zhu · 2023
Cited alongside, same era.
Improving image generation with better captions
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al · 2023
Cited alongside, same era.
Q. Dao, H. Phung, B. Nguyen, and A. Tran · 2023
Cited alongside, same era.
Masked diffusion transformer is a strong image synthesizer
S. Gao, P. Zhou, M.-M. Cheng, and S. Yan · 2023
Cited alongside, same era.
S. Li, H. Singh, and A. Grover · 2024
Closest in time.
Pointmamba: A simple state space model for point cloud analysis
D. Liang, X. Zhou, W. Xu, X. Zhu, Z. Zou, X. Ye, X. Tan, and X. Bai · 2024
Closest in time.
Swin-umamba: Mamba-based unet with imagenet-based pretraining
J. Liu, H. Yang, H.-Y. Zhou, Y. Xi, L. Yu, C. Li, Y. Liang, G. Shi, Y. Yu, S. Zhang, et al · 2024
Closest in time.
Vmamba: Visual state space model
Y. Liu, Y. Tian, Y. Zhao, H. Yu, L. Xie, Y. Wang, Q. Ye, J. Jiao, and Y. Liu · 2024
Closest in time.
U-mamba: Enhancing long-range dependency for biomedical image segmentation
J. Ma, F. Li, and B. Wang · 2024
Closest in time.
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers
N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie · 2024
Closest in time.
Can mamba learn how to learn? a comparative study on in-context learning tasks
J. Park, J. Park, Z. Xiong, N. Lee, J. Cho, S. Oymak, K. Lee, and D. Papailiopoulos · 2024
Closest in time.
Simba: Simplified mamba-based architecture for vision and multivariate time series
B. N. Patro and V. S. Agneeswaran · 2024
Closest in time.
Fast high-resolution image synthesis with latent adversarial diffusion distillation
A. Sauer, F. Boesel, T. Dockhorn, A. Blattmann, P. Esser, and R. Rombach · 2024
Closest in time.
Boosting latent diffusion with flow matching
J. Schusterbauer, M. Gui, P. Ma, N. Stracke, S. A. Baumann, V. T. Hu, and B. Ommer · 2024
Closest in time.
Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion
V. Voleti, C.-H. Yao, M. Boss, A. Letts, D. Pankratz, D. Tochilkin, C. Laforte, R. Rombach, and V. Jampani · 2024
Closest in time.
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation
Z. Wang, C. Lu, Y. Wang, F. Bao, C. Li, H. Su, and J. Zhu · 2024
Closest in time.
Diffusion models without attention
J. N. Yan, J. Gu, and A. M. Rush · 2024
Closest in time.
Fast training of diffusion models with masked transformers
H. Zheng, W. Nie, A. Vahdat, and A. Anandkumar · 2024
Closest in time.
Scalable diffusion models with state space backbone
F. Zhengcong, F. Mingyuan, Y. Changqian, and H. Jusnshi · 2024
Closest in time.
Freqmamba: Viewing mamba from a frequency perspective for image deraining
Z. Zou, H. Yu, J. Huang, and F. Zhao · 2024
Closest in time.
Depthfm: Fast monocular depth estimation with flow matching
M. Gui, J. S. Fischer, U. Prestel, P. Ma, D. Kotovenko, O. Grebenkova, S. A. Baumann, V. T. Hu, and B. Ommer · 2025
Closest in time.
Jamba: Hybrid transformer-mamba language models
B. Lenz, O. Lieber, A. Arazi, A. Bergman, A. Manevich, B. Peleg, B. Aviram, C. Almagor, C. Fridman, D. Padnos, D. Gissin, D. Jannai, D. Muhlgay, D. Zimberg, E. M. Gerber, E. Dolev, E. Krakovsky, E. Safahi, E. Schwartz, G. Cohen, G. Shachaf, H. Rozenblum, H. Bata, I. Blass, I. Magar, I. Dalmedigos, J. Osin, J. Fadlon, M. Rozman, M. Danos, M. Gokhman, M. Zusman, N. Gidron, N. Ratner, N. Gat, N. Rozen, O. Fried, O. Leshno, O. Antverg, O. Abend, O. Dagan, O. Cohavi, R. Alon, R. Belson, R. Cohen, R. Gilad, R. Glozman, S. Lev, S. Shalev-Shwartz, S. H. Meirom, T. Delbari, T. Ness, T. Asida, T. B. Gal, T. Braude, U. Pumerantz, J. Cohen, Y. Belinkov, Y. Globerson, Y. P. Levy, and Y. Shoham · 2025
Closest in time.
Point cloud mamba: Point cloud learning via state space model
T. Zhang, H. Yuan, L. Qi, J. Zhanng, Q. Zhou, S. Ji, S. Yan, and X. Li · 2025
Closest in time.