Fetching the paper…
Reading the bibliography…
State Space Model (SSM) is a mathematical model used to describe and analyze the behavior of dynamic systems.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NeurIPS , 2012, pp. 1106–1114
2012
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Dollár, “Microsoft coco: Common objects in context,” in ECCV , 2014, pp. 740–755
2014
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” arXiv , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in CVPR , 2017, pp. 77–85
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in ICCV , 2017, pp. 618–626
2017
Earlier work this paper cites.
O. Bernard, A. Lalande, C. Zotti, F. Cervenansky, X. Yang, P.-A. Heng, I. Cetin, K. Lekadir, O. Camara, M. A. Gonzalez Ballester, G. Sanroma, S. Napel, S. Petersen, G. Tziritas, E. Grinias, M. Khened, V. A. Kollerathu, G. Krishnamurthi, M.-M. Rohé, X. Pennec, M. Sermesant, F. Isensee, P. J?ger, K. H. Maier-Hein, P. M. Full, I. Wolf, S. Engelhardt, C. F. Baumgartner, L. M. Koch, J. M. Wolterink, I. I?gum, Y. Jang, Y. Hong, J. Patravali, S. Jain, O. Humbert, and P.-M. Jodoin, “Deep learning techniques for automatic mri cardiac multi-structures segmentation and diagnosis: Is the problem solved?” IEEE Transactions on Medical Imaging , vol. 37, no. 11, pp. 2514–2525, 2018
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-low-shotitation networks,” in CVPR , 2018, pp. 7132–7141
2018
Earlier work this paper cites.
Hu, Jie and Shen, Li and Sun, Gang, “Squeeze-and-excitation networks,” in CVPR , 2018, pp. 7132–7141
2018
Earlier work this paper cites.
A. Myronenko, “3d mri brain tumor segmentation using autoencoder regularization,” in BrainLes@MICCAI , 2018
2018
Earlier work this paper cites.
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun, “Unified perceptual parsing for scene understanding,” in ECCV , 2018, pp. 418–434
2018
Earlier work this paper cites.
M. Allan, A. A. Shvets, T. Kurmann, Z. Zhang, R. Duggal, Y.-H. Su, N. Rieke, I. Laina, N. Kalavakonda, S. Bodenstedt, L. C. García-Peraza, W. Li, V. I. Iglovikov, H. Luo, J. Yang, D. Stoyanov, L. Maier-Hein, S. Speidel, and M. Azizian, “2017 robotic instrument segmentation challenge,” arXiv , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
H. Liu, K. Simonyan, and Y. Yang, “Darts: Differentiable architecture search,” in ICLR , 2019, pp. 1–13
2019
Earlier work this paper cites.
M. Tan and Q. V. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in ICML , 2019, pp. 6105–6114
2019
Earlier work this paper cites.
H. X. T. S. A. A. Zhou, BoleiZhao, “Semantic understanding of scenes through the ade20k dataset,” International Journal of Computer Vision , pp. 302–321, 2019
2019
Earlier work this paper cites.
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré, “Hippo: Recurrent memory with optimal polynomial projections,” in NeurIPS , 2020, pp. 1474–1487
2020
Earlier work this paper cites.
F. Isensee, P. F. Jaeger, S. A. A. Kohl, J. Petersen, and K. Maier-Hein, “nnu-net: a self-configuring method for deep learning-based biomedical image segmentation,” Nature Methods , vol. 18, pp. 203 – 211, 2020
2020
Earlier work this paper cites.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollar, “Designing network design spaces,” in CVPR , 2020
2020
Earlier work this paper cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. J’egou, “Training data-efficient image transformers & distillation through attention,” in ICML , 2020, pp. 10 347–10 357
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid, “Vivit: A video vision transformer,” in ICCV , 2021, pp. 6836–6846
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
H. Cao, Y. Wang, J. Chen, D. Jiang, X. Zhang, Q. Tian, and M. Wang, “Swin-unet: Unet-like pure transformer for medical image segmentation,” in ECCV , 2021
2021
Earlier work this paper cites.
A. Dosovitskiy, et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021, pp. 1–16
2021
Earlier work this paper cites.
A. Gu, I. Johnson, K. Goel, K. Saab, T. Dao, A. Rudra, and C. Ré, “Combining recurrent, convolutional, and continuous-time models with linear state space layers,” in NeurIPS , 2021, pp. 572–585
2021
Earlier work this paper cites.
A. Hatamizadeh, D. Yang, H. R. Roth, and D. Xu, “Unetr: Transformers for 3d medical image segmentation,” WACV , pp. 1748–1758, 2021
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021, pp. 9992–10 002
2021
Earlier work this paper cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pvt v2: Improved baselines with pyramid vision transformer,” Computational Visual Media , pp. 415–424, 2021
2021
Earlier work this paper cites.
Y. Zhang, H. Liu, and Q. Hu, “Transfuse: Fusing transformers and cnns for medical image segmentation,” arXiv , 2021
2021
Earlier work this paper cites.
H.-Y. Zhou, J. Guo, Y. Zhang, L. Yu, L. Wang, and Y. Yu, “nnformer: Volumetric medical image segmentation via a 3d transformer,” IEEE TIP , vol. 32, pp. 4036–4045, 2021
2021
Earlier work this paper cites.
Z. L. andHanzi Mao, C. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in CVPR , 2022, pp. 11 966–11 976
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. Gu, K. Goel, and C. Ré, “Efficiently modeling long sequences with structured state spaces,” in ICLR , 2022, pp. 1–27
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Ruan, S. Xiang, M. Xie, T. Liu, and Y. Fu, “Malunet: A multi-attention and light-weight unet for skin lesion segmentation,” BIBM , pp. 1150–1156, 2022
2022
Earlier work this paper cites.
W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, and S. Yan, “Metaformer is actually what you need for vision,” in CVPR , 2022, pp. 10 819–10 829
2022
Earlier work this paper cites.
Y. Cai, H. Bian, J. Lin, H. Wang, R. Timofte, and Y. Zhang, “Retinexformer: One-stage retinex-based transformer for low-light image enhancement,” in ICCV , 2023, pp. 12 470–12 479
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. M. Islam, M. Hasan, K. S. Athrey, T. Braskich, and G. Bertasius, “Efficient movie scene detection using state-space transformers,” in CVPR , 2023, pp. 18 749–18 758
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in NeurIPS , vol. 36, 2023, pp. 34 892–34 916
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
B. N. Patro and V. S. Agneeswaran, “Mamba-360: Survey of state space models as transformer alternative for long sequence modelling: Methods, applications, and challenges,” arXiv , 2024
2024
Closest in time.
S. Peng, X. Zhu, H. Deng, Z. Lei, and L.-J. Deng, “Fusionmamba: Efficient image fusion with state space model,” arXiv , 2024
2024
Closest in time.
Z. Qian and Z. Xiao, “Smcd: High realism motion style transfer via mamba-based diffusion,” arXiv , 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Q. Wang, C. Wang, Z. Lai, and Y. Zhou, “Insectmamba: Insect pest classification with state space model,” arXiv , 2024
2024
Closest in time.
X. Wang, S. Wang, Y. Ding, Y. Li, W. Wu, Y. Rong, W. Kong, J. Huang, S. Li, H. Yang, Z. Wang, B. Jiang, C. Li, Y. Wang, Y. Tian, and J. Tang, “State space model for new-generation network alternative to transformers: A survey,” arXiv , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
P. Xiaohuan, T. Huang, and C. Xu, “Efficientvmamba: Atrous selective scan for light weight visual mamba,” arXiv , 2024
2024
Closest in time.
J. Xie, R. Liao, Z. Zhang, S. Yi, Y. Zhu, and G. Luo, “Promamba: Prompt-mamba for polyp segmentation,” arXiv , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Xu, S. Yang, Y. Wang, and B. Du, “A survey on vision mamba: Models, applications and challenges,” arXiv , 2024
2024
Closest in time.
2024
Closest in time.
G. Yang, K. Du, Z. Yang, Y. Du, and S. W. Yongping Zheng, “Cmvim: Contrastive masked vim autoencoder for 3d multi-modal representation learning for ad classification,” arXiv , 2024
2024
Closest in time.
2024
Closest in time.
Y. Yang, Z. Xing, C. Huang, and L. Zhu, “Vivim: a video vision mamba for medical video object segmentation,” arXiv , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Zhan, X. Li, H. Yuan, S. Ji, and S. Yan, “Point cloud mamba: Point cloud learning via state space model,” arXiv , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Zhang, W. Yan, K. Yan, C. Lam, Y. Qiu, P. Zheng, R. Tang, and S. Cheng, “Motion-guided dual-camera tracker for low-cost skill evaluation of gastric endoscopy,” arXiv , 2024
2024
Closest in time.
Z. Zhang, A. Liu, I. Reid, R. Hartley, B. Zhuang, and H. Tang, “Motion mamba: Efficient and long sequence motion generation with hierarchical and bidirectional selective ssm,” arXiv , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zhen, Y. Hu, and Z. Feng, “Freqmamba: Viewing mamba from a frequency perspective for image deraining,” arXiv , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, “Vision mamba: Efficient visual representation learning with bidirectional state space model,” arXiv , 2024
2024
Closest in time.