Fetching the paper…
Reading the bibliography…
The diffusion model has long been plagued by scalability and quadratic complexity issues, especially within transformer-based structures.
Anderson, B.D.: Reverse-time diffusion equation models. Stochastic Processes and their Applications (1982)
1982
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: ECCV (2014)
2014
Earlier work this paper cites.
Ronneberger, O., Fischer, P., Brox, T.: U-net: Convolutional networks for biomedical image segmentation. In: MICCAI (2015)
2015
Earlier work this paper cites.
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., Ganguli, S.: Deep unsupervised learning using nonequilibrium thermodynamics. In: ICML (2015)
2015
Earlier work this paper cites.
Newell, A., Yang, K., Deng, J.: Stacked hourglass networks for human pose estimation. In: ECCV (2016)
2016
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: NeurIPS (2017)
2017
Earlier work this paper cites.
Chen, R.T., Rubanova, Y., Bettencourt, J., Duvenaud, D.K.: Neural ordinary differential equations. NeurIPS (2018)
2018
Earlier work this paper cites.
Zhang, X., Zhou, X., Lin, M., Sun, J.: Shufflenet: An extremely efficient convolutional neural network for mobile devices. In: CVPR (2018)
2018
Earlier work this paper cites.
Child, R., Gray, S., Radford, A., Sutskever, I.: Generating long sequences with sparse transformers. arXiv (2019)
2019
Earlier work this paper cites.
Karras, T., Laine, S., Aila, T.: A style-based generator architecture for generative adversarial networks. In: CVPR (2019)
2019
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: ICLR (2019)
2019
Earlier work this paper cites.
McKenna, D.M.: Hilbert curves: Outside-in and inside-gone. Mathemaesthetics, Inc (2019)
2019
Earlier work this paper cites.
Song, Y., Ermon, S.: Generative modeling by estimating gradients of the data distribution. arXiv (2019)
2019
Earlier work this paper cites.
Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., Gelly, S.: Fvd: A new metric for video generation. ICLR Workshop (2019)
2019
Earlier work this paper cites.
Beltagy, I., Peters, M.E., Cohan, A.: Longformer: The long-document transformer. arXiv (2020)
2020
Earlier work this paper cites.
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al.: Rethinking attention with performers. arXiv (2020)
2020
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. In: NeurIPS (2020)
2020
Earlier work this paper cites.
Kitaev, N., Kaiser, Ł., Levskaya, A.: Reformer: The efficient transformer. arXiv (2020)
2020
Earlier work this paper cites.
Chefer, H., Gur, S., Wolf, L.: Transformer interpretability beyond attention visualization. In: CVPR (2021)
2021
Earlier work this paper cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. In: ICLR (2021)
2021
Earlier work this paper cites.
Esser, P., Rombach, R., Ommer, B.: Taming transformers for high-resolution image synthesis. In: CVPR (2021)
2021
Earlier work this paper cites.
Gu, A., Goel, K., Ré, C.: Efficiently modeling long sequences with structured state spaces (2021)
2021
Earlier work this paper cites.
Gu, A., Johnson, I., Goel, K., Saab, K., Dao, T., Rudra, A., Ré, C.: Combining recurrent, convolutional, and continuous-time models with linear state space layers. NeurIPS (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Kingma, D., Salimans, T., Poole, B., Ho, J.: Variational diffusion models. In: NeurIPS (2021)
2021
Earlier work this paper cites.
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: ICCV (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: ICML (2021)
2021
Earlier work this paper cites.
Skorokhodov, I., Sotnikov, G., Elhoseiny, M.: Aligning latent and image spaces to connect the unconnectable. In: ICCV (2021)
2021
Earlier work this paper cites.
Song, Y., Sohl-Dickstein, J., Kingma, D.P., Kumar, A., Ermon, S., Poole, B.: Score-based generative modeling through stochastic differential equations. In: ICLR (2021)
2021
Earlier work this paper cites.
Sun, Z., Yang, Y., Yoo, S.: Sparse attention with learning to hash. In: ICLR (2021)
2021
Earlier work this paper cites.
Xia, W., Yang, Y., Xue, J.H., Wu, B.: Tedigan: Text-guided diverse face image generation and manipulation. In: CVPR (2021)
2021
Earlier work this paper cites.
Albergo, M.S., Vanden-Eijnden, E.: Building normalizing flows with stochastic interpolants. arXiv (2022)
2022
Earlier work this paper cites.
Ben-Hamu, H., Cohen, S., Bose, J., Amos, B., Grover, A., Nickel, M., Chen, R.T., Lipman, Y.: Matching normalizing flows and probability paths on manifolds. In: ICML (2022)
2022
Earlier work this paper cites.
Dao, T., Fu, D., Ermon, S., Rudra, A., Ré, C.: Flashattention: Fast and memory-efficient exact attention with io-awareness. NeurIPS (2022)
2022
Earlier work this paper cites.
Fu, D.Y., Dao, T., Saab, K.K., Thomas, A.W., Rudra, A., Ré, C.: Hungry hungry hippos: Towards language modeling with state space models. arXiv (2022)
2022
Earlier work this paper cites.
Gu, A., Goel, K., Gupta, A., Ré, C.: On the parameterization and initialization of diagonal state space models. NeurIPS (2022)
2022
Earlier work this paper cites.
Gupta, A., Gu, A., Berant, J.: Diagonal state spaces are as effective as structured state spaces. NeurIPS (2022)
2022
Earlier work this paper cites.
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., Cohen-Or, D.: Prompt-to-prompt image editing with cross attention control. arXiv (2022)
2022
Earlier work this paper cites.
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., Fleet, D.J.: Video diffusion models. In: ARXIV (2022)
2022
Earlier work this paper cites.
Karras, T., Aittala, M., Aila, T., Laine, S.: Elucidating the design space of diffusion-based generative models. In: NeurIPS (2022)
2022
Earlier work this paper cites.
Liu, G.H., Chen, T., So, O., Theodorou, E.: Deep generalized schrödinger bridge. NeurIPS (2022)
2022
Earlier work this paper cites.
Liu, X., Gong, C., Liu, Q.: Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv (2022)
2022
Earlier work this paper cites.
Nguyen, E., Goel, K., Gu, A., Downs, G., Shah, P., Dao, T., Baccus, S., Ré, C.: S4nd: Modeling images and videos as multidimensional signals with state spaces. NeurIPS (2022)
2022
Cited alongside, same era.
Peebles, W., Xie, S.: Scalable diffusion models with transformers. arXiv (2022)
2022
Cited alongside, same era.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR (2022)
2022
Cited alongside, same era.
Smith, J.T., Warrington, A., Linderman, S.W.: Simplified state space layers for sequence modeling. arXiv (2022)
2022
Cited alongside, same era.
Tang, R., Liu, L., Pandey, A., Jiang, Z., Yang, G., Kumar, K., Stenetorp, P., Lin, J., Ture, F.: What the daam: Interpreting stable diffusion using cross attention. arXiv (2022)
2022
Cited alongside, same era.
2024
Closest in time.
Gong, H., Kang, L., Wang, Y., Wan, X., Li, H.: nnmamba: 3d biomedical image segmentation, classification and landmark detection with state space model. arXiv (2024)
2024
Closest in time.
Gu, A., Dao, T.: Mamba: Linear-time sequence modeling with selective state spaces. CoLM (2024)
2024
Closest in time.
2024
Closest in time.
Guo, H., Li, J., Dai, T., Ouyang, Z., Ren, X., Xia, S.T.: Mambair: A simple baseline for image restoration with state-space model. arXiv (2024)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, J., Yan, J.N., Gu, A., Rush, A.M.: Pretraining without attention. arXiv (2022)
2022
Cited alongside, same era.
Agarwal, N., Suo, D., Chen, X., Hazan, E.: Spectral state space models. arXiv (2023)
2023
Cited alongside, same era.
Albergo, M.S., Boffi, N.M., Vanden-Eijnden, E.: Stochastic interpolants: A unifying framework for flows and diffusions. arXiv (2023)
2023
Cited alongside, same era.
Bao, F., Li, C., Cao, Y., Zhu, J.: All are worth words: a vit backbone for score-based diffusion models. CVPR (2023)
2023
Cited alongside, same era.
Bao, F., Nie, S., Xue, K., Li, C., Pu, S., Wang, Y., Yue, G., Cao, Y., Su, H., Zhu, J.: One transformer fits all distributions in multi-modal diffusion at scale. arXiv (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Chen, S., Xu, M., Ren, J., Cong, Y., He, S., Xie, Y., Sinha, A., Luo, P., Xiang, T., Perez-Rua, J.M.: Gentron: Delving deep into diffusion transformers for image and video generation. arXiv (2023)
2023
Cited alongside, same era.
2024
Closest in time.
He, W., Han, K., Tang, Y., Wang, C., Yang, Y., Guo, T., Wang, Y.: Densemamba: State space models with dense hidden connection for efficient large language models. arXiv (2024)
2024
Closest in time.
He, X., Cao, K., Yan, K., Li, R., Xie, C., Zhang, J., Zhou, M.: Pan-mamba: Effective pan-sharpening with state space model. arXiv (2024)
2024
Closest in time.
Hu, V.T., Wu, D., Asano, Y., Mettes, P., Fernando, B., Ommer, B., Snoek, C.: Flow matching for conditional text generation in a few sampling steps pp. 380–392 (2024)
2024
Closest in time.
Huang, Z., Zhou, P., Yan, S., Lin, L.: Scalelong: Towards more stable training of diffusion model via scaling network long skip connection. NeurIPS (2024)
2024
Closest in time.
Li, K., Li, X., Wang, Y., He, Y., Wang, Y., Wang, L., Qiao, Y.: Videomamba: State space model for efficient video understanding. ECCV (2024)
2024
Closest in time.
Li, S., Singh, H., Grover, A.: Mamba-nd: Selective state space modeling for multi-dimensional data. arXiv (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Lin, B., Jiang, W., Chen, P., Zhang, Y., Liu, S., Chen, Y.C.: Mtmamba: Enhancing multi-task dense scene understanding by mamba-based decoders. ECCV (2024)
2024
Closest in time.
Liu, J., Yang, H., Zhou, H.Y., Xi, Y., Yu, L., Yu, Y., Liang, Y., Shi, G., Zhang, S., Zheng, H., et al.: Swin-umamba: Mamba-based unet with imagenet-based pretraining. arXiv (2024)
2024
Closest in time.
Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., Liu, Y.: Vmamba: Visual state space model. arXiv (2024)
2024
Closest in time.
Ma, J., Li, F., Wang, B.: U-mamba: Enhancing long-range dependency for biomedical image segmentation. arXiv (2024)
2024
Closest in time.
Ma, N., Goldstein, M., Albergo, M.S., Boffi, N.M., Vanden-Eijnden, E., Xie, S.: Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers. ECCV (2024)
2024
Closest in time.
OpenAI: Sora: Creating video from text (2024), https://openai.com/sora
2024
Closest in time.
Park, J., Kim, H.S., Ko, K., Kim, M., Kim, C.: Videomamba: Spatio-temporal selective state space model. ECCV (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Ruan, J., Xiang, S.: Vm-unet: Vision mamba unet for medical image segmentation. arXiv (2024)
2024
Closest in time.
Tikochinski, R., Goldstein, A., Meiri, Y., Hasson, U., Reichart, R.: An incremental large language model for long text processing in the brain (2024)
2024
Closest in time.
Wang, C., Tsepa, O., Ma, J., Wang, B.: Graph-mamba: Towards long-range graph sequence modeling with selective state spaces. arXiv (2024)
2024
Closest in time.
Wang, J., Gangavarapu, T., Yan, J.N., Rush, A.M.: Mambabyte: Token-free selective state space model. arXiv (2024)
2024
Closest in time.
Wang, S., Xue, B.: State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory. NeurIPS (2024)
2024
Closest in time.
Wang, X., Wang, S., Ding, Y., Li, Y., Wu, W., Rong, Y., Kong, W., Huang, J., Li, S., Yang, H., Wang, Z., Jiang, B., Li, C., Wang, Y., Tian, Y., Tang, J.: State space model for new-generation network alternative to transformers: A survey (2024)
2024
Closest in time.
2024
Closest in time.
Wang, Z., Ma, C.: Semi-mamba-unet: Pixel-level contrastive cross-supervised visual mamba-based unet for semi-supervised medical image segmentation. arXiv (2024)
2024
Closest in time.
Wang, Z., Zheng, J.Q., Zhang, Y., Cui, G., Li, L.: Mamba-unet: Unet-like pure visual mamba for medical image segmentation. arXiv (2024)
2024
Closest in time.
Xing, Z., Ye, T., Yang, Y., Liu, G., Zhu, L.: Segmamba: Long-range sequential modeling mamba for 3d medical image segmentation. arXiv (2024)
2024
Closest in time.
Yang, S., Wang, B., Shen, Y., Panda, R., Kim, Y.: Gated linear attention transformers with hardware-efficient training. ICML (2024)
2024
Closest in time.
Yang, S., Zhang, Y.: Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism (Jan 2024), https://github.com/sustcsonglin/flash-linear-attention
2024
Closest in time.
Yang, Y., Xing, Z., Zhu, L.: Vivim: a video vision mamba for medical video object segmentation. arXiv (2024)
2024
Closest in time.
Zhang, T., Li, X., Yuan, H., Ji, S., Yan, S.: Point could mamba: Point cloud learning via state space model. arXiv (2024)
2024
Closest in time.
Zhang, Z., Liu, A., Reid, I., Hartley, R., Zhuang, B., Tang, H.: Motion mamba: Efficient and long sequence motion generation with hierarchical and bidirectional selective ssm. ECCV (2024)
2024
Closest in time.
Zhang, Z., Liu, A., Reid, I., Hartley, R., Zhuang, B., Tang, H.: Motion mamba: Efficient and long sequence motion generation with hierarchical and bidirectional selective ssm. arXiv (2024)
2024
Closest in time.
Zheng, Z., Wu, C.: U-shaped vision mamba for single image dehazing. arXiv (2024)
2024
Closest in time.
Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., Wang, X.: Vision mamba: Efficient visual representation learning with bidirectional state space model. ICML (2024)
2024
Closest in time.
zhuzilin: Ring flash attention. https://github.com/zhuzilin/ring-flash-attention (2024)
2024
Closest in time.