Fetching the paper…
Reading the bibliography…
Structured State Space Models (SSMs) have become a prominent class of sequence models, developed against two long-standing difficulties: the sequential computation and gradient propagation limits of Recurrent Neural Networks (RNNs), and the quadratic time and memory cost of self-attention in Transformers.
Coupled mamba: Enhanced multimodal fusion with coupled state space model
Wenbing Li, Hang Zhou, Junqing Yu, Zikai Song, and Wei Yang · 1910
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman · 1960
Earlier work this paper cites.
Linear Systems , volume 156
Thomas Kailath · 1980
Earlier work this paper cites.
System Identification: Theory for the User
Lennart Ljung · 1987
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S. Denker, Donnie Henderson, Richard E. Howard, Wayne Hubbard, and Lawrence D. Jackel · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Time Series Analysis
James D. Hamilton · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Serial order: A parallel distributed processing approach
Michael I. Jordan · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Analysis and control of nonlinear process systems
Katalin M. Hangos, József Bokor, and Gábor Szederkényi · 2006
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Neural ordinary differential equations
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David K. Duvenaud · 2018
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, et al · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Using cnn for facial expression recognition: A study of the effects of kernel size and number of filters on accuracy
Abhinav Agrawal and Namita Mittal · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
HiPPO: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity, 2020
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Earlier work this paper cites.
Spectral normalisation for deep reinforcement learning: An optimisation perspective
Florin Gogianu, Tudor Berariu, Mihaela C. Rosca, Claudia Clopath, Lucian Busoniu, and Razvan Pascanu · 2021
Earlier work this paper cites.
Combining recurrent, convolutional, and continuous-time models with linear state-space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Earlier work this paper cites.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Earlier work this paper cites.
Temporal fusion transformers for interpretable multi-horizon time series forecasting
Bryan Lim, Sercan Ö Arık, Nicolas Loeff, and Tomas Pfister · 2021
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Cited alongside, same era.
A novel representation learning for dynamic graphs based on graph convolutional networks
Chao Gao, Junyou Zhu, Fan Zhang, Zhen Wang, and Xuelong Li · 2022
Cited alongside, same era.
On the parameterization and initialization of diagonal state space models
Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré · 2022
Cited alongside, same era.
Diagonal state spaces are as effective as structured state spaces
Ankit Gupta, Albert Gu, and Jonathan Berant · 2022
Cited alongside, same era.
S4ND: Modeling images and videos as multidimensional signals with state spaces
Eric Nguyen, Karan Goel, Albert Gu, Gordon W. Downs, Preey Shah, Tri Dao, Stephen Baccus, and Christopher Ré · 2022
Cited alongside, same era.
Zamba: A compact 7b SSM hybrid model
P. Glorioso, Q. Anthony, Y. Tokpanov, J. Whittington, J. Pilault, A. Ibrahim, and B. Millidge · 2024
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2024
Later among the works it cites.
MambaAD: Exploring state space models for multi-class unsupervised anomaly detection
Haoyang He, Yuhu Bai, Jiangning Zhang, Qingdong He, Hongxu Chen, Zhenye Gan, Chengjie Wang, Xiangtai Li, Guanzhong Tian, and Lei Xie · 2024
Later among the works it cites.
Computation-efficient era: A comprehensive survey of state space models in medical image analysis, 2024
Moein Heidari, Sina Ghorbani Kolahi, Sanaz Karimijafarbigloo, Bobby Azad, Afshin Bozorgpour, Soheila Hatami, Reza Azad, Ali Diba, Ulas Bagci, Dorit Merhof, and Ilker Hacihaliloglu · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2022
Cited alongside, same era.
Efficient long sequence modeling via state space augmented transformer, 2022
Simiao Zuo, Xiaodong Liu, Jian Jiao, Denis Charles, Eren Manavoglu, Tuo Zhao, and Jianfeng Gao · 2022
Cited alongside, same era.
A comparison of lstm and gru networks for learning symbolic sequences
Roberto Cahuantzi, Xinye Chen, and Stefan Güttel · 2023
Cited alongside, same era.
Hungry hungry hippos: Towards language modeling with state space models
Daniel Y. Fu, Tri Dao, Khaled K. Saab, Armin W. Thomas, Atri Rudra, and Christopher Ré · 2023
Cited alongside, same era.
How to train your HiPPO: State space models with generalized orthogonal basis projections
Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Ré · 2023
Cited alongside, same era.
Liquid structural state-space models
Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, and Daniela Rus · 2023
Cited alongside, same era.
Mega: Moving average equipped gated attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer · 2023
Cited alongside, same era.
Jamba Team · 2024
Later among the works it cites.
DGMamba: Domain generalization via generalized state space model
Shaocong Long, Qianyu Zhou, Xiangtai Li, Xuequan Lu, Chenhao Ying, Yuan Luo, Lizhuang Ma, and Shuicheng Yan · 2024
Later among the works it cites.
U-mamba: Enhancing long-range dependency for biomedical image segmentation
Jun Ma, Feifei Li, and Bo Wang · 2024
Later among the works it cites.
Exploring the capability of mamba in speech applications
Koichi Miyazaki, Yoshiki Masuyama, and Masato Murata · 2024
Later among the works it cites.
Simba: Simplified mamba-based architecture for vision and multivariate time series, 2024
Badri N. Patro and Vijay S. Agneeswaran · 2024
Later among the works it cites.
VL-Mamba: Exploring state space models for multimodal learning
Yanyuan Qiao, Zheng Yu, Zijia Zhao, Sihan Chen, Mingzhen Sun, Longteng Guo, Qi Wu, and Jing Liu · 2024
Later among the works it cites.
HGRN2: Gated linear RNNs with state expansion
Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong · 2024
Later among the works it cites.
StableSSM: Alleviating the curse of memory in state-space models through stable reparameterization
Shida Wang and Qianxiao Li · 2024
Later among the works it cites.
FusionMamba: dynamic feature enhancement for multimodal image fusion with Mamba
Xinyu Xie, Yawen Cui, Tao Tan, Xubin Zheng, and Zitong Yu · 2024
Later among the works it cites.
Visual mamba: A survey and new outlooks
Rui Xu, Shu Yang, Yihui Wang, Yu Cai, Bo Du, and Hao Chen · 2024
Later among the works it cites.
Linear recurrent units for sequential recommendation
Zhenrui Yue, Yueqi Wang, Zhankui He, Huimin Zeng, Julian McAuley, and Dong Wang · 2024
Later among the works it cites.
Vision mamba in remote sensing: A comprehensive survey of techniques, applications and outlook
Muyi Bao, Shuchang Lyu, Zhaoyang Xu, Huiyu Zhou, Jinchang Ren, Shiming Xiang, Xiangtai Li, and Guangliang Cheng · 2025
Closest in time.
Zifeng Ding, Yifeng Li, Yuan He, Antonio Norelli, Jingcheng Wu, Volker Tresp, Michael M. Bronstein, and Yunpu Ma · 2025
Closest in time.
Hymba: A hybrid-head architecture for small language models
X. Dong, Y. Fu, S. Diao, W. Byeon, Z. Chen, A. Mahabaleshwarkar, S.-Y. Liu, M. Van Keirsbilck, M.-H. Chen, Y. Suhara, Y. C. Lin, J. Kautz, and P. Molchanov · 2025
Closest in time.
Jamba: Hybrid transformer-mamba language models
Barak Lenz, Opher Lieber, Alan Arazi, Amir Bergman, Avshalom Manevich, Barak Peleg, Ben Aviram, Chen Almagor, Clara Fridman, Dan Padnos, Daniel Gissin, Daniel Jannai, Dor Muhlgay, Dor Zimberg, Edden M. Gerber, Elad Dolev, Eran Krakovsky, Erez Safahi, Erez Schwartz, Gal Cohen, Gal Shachaf, Haim Rozenblum, Hofit Bata, Ido Blass, Inbal Magar, Itay Dalmedigos, Jhonathan Osin, Julie Fadlon, Maria Rozman, Matan Danos, Michael Gokhman, Mor Zusman, Naama Gidron, Nir Ratner, Noam Gat, Noam Rozen, Oded Fried, Ohad Leshno, Omer Antverg, Omri Abend, Or Dagan, Orit Cohavi, Raz Alon, Ro’i Belson, Roi Cohen, Rom Gilad, Roman Glozman, Shahar Lev, Shai Shalev-Shwartz, Shaked Haim Meirom, Tal Delbari, Tal Ness, Tomer Asida, Tom Ben Gal, Tom Braude, Uriya Pumerantz, Josh Cohen, Yonatan Belinkov, Yuval Globerson, Yuval Peleg Levy, and Yoav Shoham · 2025
Closest in time.
A survey of multimodal fake news detection: A cross-modal interaction perspective
Xianghua Li, Jiao Qiao, Shu Yin, Lianwei Wu, Chao Gao, Zhen Wang, and Xuelong Li · 2025
Closest in time.
A survey of RWKV
Zhiyuan Li, Tingyu Xia, Yi Chang, and Yuan Wu · 2025
Closest in time.
Mamba-360: Survey of state space models as transformer alternative for long sequence modelling: Methods, applications, and challenges
Badri Narayana Patro and Vijay Srinivas Agneeswaran · 2025
Closest in time.
Samba: Simple hybrid state space models for efficient unlimited context language modeling
L. Ren, Y. Liu, Y. Lu, Y. Shen, C. Liang, and W. Chen · 2025
Closest in time.
A comprehensive survey on Mamba: Architectures, challenges, and opportunities
Abdus Salam, Rasel Mahmud, Tohedul Islam, Saddam Mukta, and Swakkhar Shatabda · 2025
Closest in time.
Crash severity analysis of child bicyclists using Arm-Net and MambaNet
Shriyank Somvanshi, Rohit Chakraborty, Subasish Das, and Anandi K. Dutta · 2025
Closest in time.
Applying tabular deep learning models to estimate crash injury types of young motorcyclists
Shriyank Somvanshi, Anannya Ghosh Tusti, Rohit Chakraborty, and Subasish Das · 2025
Closest in time.
Multi-scale graph contrastive learning for community detection in dynamic graphs
Min Teng, Chao Gao, Xianghua Li, Zhen Wang, Kefeng Fan, and Vladimir Nekorkin · 2025
Closest in time.
Multilingual state space models for structured question answering in Indic languages
Arpita Vats, Rahul Raja, Mrinal Mathur, Aman Chadha, and Vinija Jain · 2025
Closest in time.
Point cloud mamba: Point cloud learning via state space model
Tao Zhang, Haobo Yuan, Lu Qi, Jiangning Zhang, Qianyu Zhou, Shunping Ji, Shuicheng Yan, and Xiangtai Li · 2025
Closest in time.
Haohao Qu, Liangbo Ning, Rui An, Wenqi Fan, Tyler Derr, Hui Liu, Xin Xu, and Qing Li · 2026
Closest in time.