Fetching the paper…
Reading the bibliography…
In recent years, State Space Models (SSMs) with efficient hardware-aware designs, known as the Mamba deep learning models, have made significant progress in modeling long sequences such as language understanding.
“Accurate, large minibatch SGD: training imagenet in 1 hour,”
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He, · 2017
Earlier work this paper cites.
“Transformers are rnns: Fast autoregressive transformers with linear attention,”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret, · 2020
Earlier work this paper cites.
“Designing betwork design spaces,”
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár, · 2020
Earlier work this paper cites.
“Training data-efficient image transformers & distillation through attention,”
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou, · 2021
Earlier work this paper cites.
“Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer,”
Sachin Mehta and Mohammad Rastegari, · 2021
Earlier work this paper cites.
“Swin transformer: Hierarchical vision transformer using shifted windows,”
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo, · 2021
Earlier work this paper cites.
“Benchmarking detection transfer learning with vision transformers,”
Yanghao Li, Saining Xie, Xinlei Chen, Piotr Dollár, Kaiming He, and Ross B. Girshick, · 2021
Earlier work this paper cites.
“Efficientformer: Vision transformers at mobilenet speed,”
Yanyu Li, Geng Yuan, Yang Wen, Ju Hu, Georgios Evangelidis, Sergey Tulyakov, Yanzhi Wang, and Jian Ren, · 2022
Cited alongside, same era.
“Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer,”
Sachin Mehta and Mohammad Rastegari, · 2022
Cited alongside, same era.
“Pvt v2: Improved baselines with pyramid vision transformer,”
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao, · 2022
Cited alongside, same era.
“Lite vision transformer with enhanced self-attention,”
Chenglin Yang, Yilin Wang, Jianming Zhang, He Zhang, Zijun Wei, Zhe Lin, and Alan Yuille, · 2022
Cited alongside, same era.
“A convnet for the 2020s,”
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie, · 2022
Cited alongside, same era.
“Mamba: Linear-time sequence modeling with selective state spaces,”
“Vision mamba: Efficient visual representation learning with bidirectional state space model,”
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang, · 2024
Closest in time.
“Vmamba: Visual state space model,”
Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, and Yunfan Liu, · 2024
Closest in time.
“Localmamba: Visual state space model with windowed selective scan,”
Tao Huang, Xiaohuan Pei, Shan You, Fei Wang, Chen Qian, and Chang Xu, · 2024
Closest in time.
“Efficientvmamba: Atrous selective scan for light weight visual mamba,”
Xiaohuan Pei, Tao Huang, and Chang Xu, · 2024
Closest in time.
“Plainmamba: Improving non-hierarchical mamba in visual recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Albert Gu and Tri Dao, · 2023
Cited alongside, same era.
“Long range language modeling via gated state spaces,”
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur, · 2023
Cited alongside, same era.
Chenhongyi Yang, Zehui Chen, Miguel Espinosa, Linus Ericsson, Zhenyu Wang, Jiaming Liu, and Elliot J. Crowley, · 2024
Closest in time.
Qinfeng Zhu, Yuan Fang, Yuanzhi Cai, Cheng Chen, and Lei Fan, · 2024
Closest in time.