Fetching the paper…
Reading the bibliography…
As one of the most representative DL techniques, Transformer architecture has empowered numerous advanced models, especially the large language models (LLMs) that comprise billions of parameters, becoming a cornerstone in deep learning.
Freqmamba: Viewing mamba from a frequency perspective for image deraining. In Proceedings of the 32nd ACM international conference on multimedia . 1905–1914
Zhen Zou, Hu Yu, Jie Huang, and Feng Zhao. 2024 · 1914
Earlier work this paper cites.
ST-mamba: spatial-temporal mamba for traffic flow estimation recovery using limited data. In 2024 IEEE/CIC International Conference on Communications in China (ICCC) . IEEE, 1928–1933
Doncheng Yuan, Jianzhe Xue, Jinshan Su, Wenchao Xu, and Haibo Zhou. 2024b · 1933
Earlier work this paper cites.
Mamba in Speech: Towards an Alternative to Self-Attention
Xiangyu Zhang, Qiquan Zhang, Hexin Liu, Tianyi Xiao, Xinyuan Qian, Beena Ahmed, Eliathamby Ambikairajah, Haizhou Li, and Julien Epps. 2025c · 1948
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
RE Kalman. 1960 · 1960
Earlier work this paper cites.
Back to the future: Towards explainable temporal reasoning with large language models. In Proceedings of the ACM on Web Conference 2024 . 1963–1974
Chenhan Yuan, Qianqian Xie, Jimin Huang, and Sophia Ananiadou. 2024a · 1974
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991 · 1991
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal. 1997 · 1997
Earlier work this paper cites.
Parallel prefix sum (scan) with CUDA
Mark Harris, Shubhabrata Sengupta, and John D Owens. 2007 · 2007
Earlier work this paper cites.
Comparison between first-order hold with zero-order hold in discretization of input-delay nonlinear systems. In 2007 International Conference on Control, Automation and Systems . IEEE, 2892–2896
Zheng Zhang and Kil To Chong. 2007 · 2007
Earlier work this paper cites.
Adversarial machine learning. In Proceedings of the 4th ACM workshop on Security and artificial intelligence . 43–58
Ling Huang, Anthony D Joseph, Blaine Nelson, Benjamin IP Rubinstein, and J Doug Tygar. 2011 · 2011
Earlier work this paper cites.
Learning auto-regressive models from sequence and non-sequence data
Tzu-Kuo Huang and Jeff Schneider. 2011 · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks. In Proceedings of the 28th international conference on machine learning (ICML-11) . 1017–1024
Ilya Sutskever, James Martens, and Geoffrey E Hinton. 2011 · 2011
Earlier work this paper cites.
Long short-term memory
Alex Graves and Alex Graves. 2012 · 2012
Earlier work this paper cites.
Training and analyzing deep recurrent neural networks. In Proceedings of the 27th International Conference on Neural Information Processing Systems-Volume 1 . 190–198
Michiel Hermans and Benjamin Schrauwen. 2013 · 2013
Earlier work this paper cites.
Convolutional neural networks for speech recognition
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang, Li Deng, Gerald Penn, and Dong Yu. 2014 · 2014
Earlier work this paper cites.
A recursive recurrent neural network for statistical machine translation. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1491–1500
Shujie Liu, Nan Yang, Mu Li, and Ming Zhou. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18 . Springer, 234–241
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Small molecules, big targets: drug discovery faces the protein–protein interaction challenge
Duncan E Scott, Andrew R Bayly, Chris Abell, and John Skidmore. 2016 · 2016
Earlier work this paper cites.
State-frequency memory recurrent neural networks. In International Conference on Machine Learning . PMLR, 1568–1577
Hao Hu and Guo-Jun Qi. 2017 · 2017
Earlier work this paper cites.
Lattice-based recurrent neural network encoders for neural machine translation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 31
Jinsong Su, Zhixing Tan, Deyi Xiong, Rongrong Ji, Xiaodong Shi, and Yang Liu. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Explainable artificial intelligence: A survey. In 2018 41st International convention on information and communication technology, electronics and microelectronics (MIPRO) . IEEE, 0210–0215
Filip Karlo Došilović, Mario Brčić, and Nikica Hlupić. 2018 · 2018
Earlier work this paper cites.
Deep modeling of social relations for recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32
Wenqi Fan, Qing Li, and Min Cheng. 2018 · 2018
Earlier work this paper cites.
Measuring catastrophic forgetting in neural networks. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32
Ronald Kemker, Marc McClure, Angelina Abitino, Tyler Hayes, and Christopher Kanan. 2018 · 2018
Earlier work this paper cites.
Micro behaviors: A new perspective in e-commerce recommender systems. In Proceedings of the eleventh ACM international conference on web search and data mining . 727–735
Meizi Zhou, Zhuoye Ding, Jiliang Tang, and Dawei Yin. 2018 · 2018
Earlier work this paper cites.
BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer. In Proceedings of the 28th ACM international conference on information and knowledge management . 1441–1450
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang. 2019 · 2019
Earlier work this paper cites.
Session-based recommendation with graph neural networks. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 346–353
Shu Wu, Yuyuan Tang, Yanqiao Zhu, Liang Wang, Xing Xie, and Tieniu Tan. 2019 · 2019
Earlier work this paper cites.
Measuring and relieving the over-smoothing problem for graph neural networks from the topological view. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 3438–3445
Deli Chen, Yankai Lin, Wei Li, Peng Li, Jie Zhou, and Xu Sun. 2020 · 2020
Earlier work this paper cites.
Human activity recognition using magnetic induction-based motion signals and deep recurrent neural networks
Negar Golestani and Mahta Moghaddam. 2020 · 2020
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré. 2020 · 2020
Earlier work this paper cites.
Deep learning for 3d point clouds: A survey
Yulan Guo, Hanyun Wang, Qingyong Hu, Hao Liu, Li Liu, and Mohammed Bennamoun. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
SkipGNN: predicting molecular interactions with skip-graph networks
Kexin Huang, Cao Xiao, Lucas M Glass, Marinka Zitnik, and Jimeng Sun. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
An semg-controlled 3d game for rehabilitation therapies: Real-time time hand gesture recognition using deep learning techniques
Nadia Nasri, Sergio Orts-Escolano, and Miguel Cazorla. 2020 · 2020
Earlier work this paper cites.
Vivit: A video vision transformer. In Proceedings of the IEEE/CVF international conference on computer vision . 6836–6846
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021 · 2021
Earlier work this paper cites.
Speaker recognition based on deep learning: An overview
Zhongxin Bai and Xiao-Lei Zhang. 2021 · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré. 2021 · 2021
Earlier work this paper cites.
Compacter: Efficient low-rank hypercomplex adapter layers
Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021 · 2021
Earlier work this paper cites.
Soft: Softmax-free transformer with linear complexity
Jiachen Lu, Jinghan Yao, Junge Zhang, Xiatian Zhu, Hang Xu, Weiguo Gao, Chunjing Xu, Tao Xiang, and Li Zhang. 2021 · 2021
Earlier work this paper cites.
Automatic speech recognition: a survey
Mishaim Malik, Muhammad Kamran Malik, Khawar Mehmood, and Imran Makhdoom. 2021 · 2021
Earlier work this paper cites.
A survey on document-level neural machine translation: Methods and evaluation
Sameen Maruf, Fahimeh Saleh, and Gholamreza Haffari. 2021 · 2021
Earlier work this paper cites.
Efficient attention: Attention with linear complexities. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 3531–3539
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li. 2021 · 2021
Earlier work this paper cites.
Motion representations for articulated animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13653–13662
Aliaksandr Siarohin, Oliver J Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov. 2021 · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention. In International conference on machine learning . PMLR, 10347–10357
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. 2021 · 2021
Earlier work this paper cites.
Sparse graph attention networks
Yang Ye and Shihao Ji. 2021 · 2021
Earlier work this paper cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence , Vol. 35. 11106–11115
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021 · 2021
Earlier work this paper cites.
Long-short transformer: Efficient transformers for language and vision
Chen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi, Tom Goldstein, Anima Anandkumar, and Bryan Catanzaro. 2021 · 2021
Earlier work this paper cites.
Graph Trend Filtering Networks for Recommendation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 112–121
Wenqi Fan, Xiaorui Liu, Wei Jin, Xiangyu Zhao, Jiliang Tang, and Qing Li. 2022 · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2022 · 2022
Earlier work this paper cites.
On the parameterization and initialization of diagonal state space models
Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré. 2022b · 2022
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations . 12513–12525
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Preventing catastrophic forgetting in continual learning of new natural language tasks. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 3137–3145
Sudipta Kar, Giuseppe Castellucci, Simone Filice, Shervin Malmasi, and Oleg Rokhlenko. 2022 · 2022
Earlier work this paper cites.
An empirical survey on long document summarization: Datasets, models, and metrics
Huan Yee Koh, Jiaxin Ju, Ming Liu, and Shirui Pan. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting
Tomasz Korbak, Hady Elsahar, Germán Kruszewski, and Marc Dymetman. 2022b · 2022
Earlier work this paper cites.
Deep learning based assistive technology on audio visual speech recognition for hearing impaired
L Ashok Kumar, D Karthika Renuka, S Lovelyn Rose, I Made Wartana, et al · 2022
Earlier work this paper cites.
Ds-transunet: Dual swin transformer u-net for medical image segmentation
Ailiang Lin, Bingzhi Chen, Jiayu Xu, Zheng Zhang, Guangming Lu, and David Zhang. 2022 · 2022
Earlier work this paper cites.
Trustworthy ai: A computational perspective
Haochen Liu, Yiqi Wang, Wenqi Fan, Xiaorui Liu, Yaxin Li, Shaili Jain, Yunhao Liu, Anil Jain, and Jiliang Tang. 2022b · 2022
Earlier work this paper cites.
Delivering trustworthy AI through formal XAI. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 12342–12350
Joao Marques-Silva and Alexey Ignatiev. 2022 · 2022
Earlier work this paper cites.
Zero-order hold discretization of general state space systems with input delay
Georgia Pechlivanidou and Nicholas Karampetakis. 2022 · 2022
Earlier work this paper cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2022 · 2022
Earlier work this paper cites.
Continual learning with lifelong vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 171–181
Zhen Wang, Liu Liu, Yiqun Duan, Yajing Kong, and Dacheng Tao. 2022 · 2022
Earlier work this paper cites.
Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 19313–19322
Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, and Jiwen Lu. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Fairly adaptive negative sampling for recommendations. In Proceedings of the ACM Web Conference 2023 . 3723–3733
Xiao Chen, Wenqi Fan, Jingfan Chen, Haochen Liu, Zitao Liu, Zhaoxiang Zhang, and Qing Li. 2023 · 2023
Earlier work this paper cites.
Towards next-generation intelligent assistants leveraging llm techniques. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 5792–5793
Xin Luna Dong, Seungwhan Moon, Yifan Ethan Xu, Kshitiz Malik, and Zhou Yu. 2023 · 2023
Earlier work this paper cites.
Adversarial Attacks for Black-Box Recommender Systems Via Copying Transferable Cross-Domain User Profiles
Wenqi Fan, Xiangyu Zhao, Qing Li, Tyler Derr, Yao Ma, Hui Liu, Jianping Wang, and Jiliang Tang. 2023 · 2023
Earlier work this paper cites.
Simple hardware-efficient long convolutions for sequence modeling. In International Conference on Machine Learning . PMLR, 10373–10391
Daniel Y Fu, Elliot L Epstein, Eric Nguyen, Armin W Thomas, Michael Zhang, Tri Dao, Atri Rudra, and Christopher Ré. 2023 · 2023
Earlier work this paper cites.
Seat: stable and explainable attention. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 12907–12915
Lijie Hu, Yixin Liu, Ninghao Liu, Mengdi Huai, Lichao Sun, and Di Wang. 2023 · 2023
Earlier work this paper cites.
Mamba-Chat
Justus Mattern and Konstantin Hohr. 2023 · 2023
Earlier work this paper cites.
Hyena hierarchy: Towards larger convolutional language models. In International Conference on Machine Learning . PMLR, 28043–28078
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré. 2023 · 2023
Earlier work this paper cites.
How will generative AI disrupt data science in drug discovery?
Jean-Philippe Vert. 2023 · 2023
Earlier work this paper cites.
3D dynamic image modeling based on machine learning in film and television animation
Yuwei Wang et al · 2023
Earlier work this paper cites.
TF-GridNet: Making time-frequency domain models great again for monaural speaker separation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
Zhong-Qiu Wang, Samuele Cornell, Shukjae Choi, Younglo Lee, Byeong-Yeol Kim, and Shinji Watanabe. 2023a · 2023
Earlier work this paper cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Earlier work this paper cites.
A unified framework for U-Net design and analysis
Christopher Williams, Fabian Falck, George Deligiannidis, Chris C Holmes, Arnaud Doucet, and Saifuddin Syed. 2023 · 2023
Cited alongside, same era.
Multimodal large language models: A survey. In 2023 IEEE International Conference on Big Data (BigData) . IEEE, 2247–2256
Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, and S Yu Philip. 2023 · 2023
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models. In International Conference on Machine Learning . PMLR, 38087–38099
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023 · 2023
Cited alongside, same era.
A survey of controllable text generation using transformer-based pre-trained language models
Hanqing Zhang, Haolin Song, Shaoyu Li, Ming Zhou, and Dawei Song. 2023 · 2023
Cited alongside, same era.
Caduceus: Bi-directional equivariant long-range dna sequence modeling
Yair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao, Albert Gu, and Volodymyr Kuleshov. 2024 · 2024
Closest in time.
Serpent: Scalable and Efficient Image Restoration via Multi-scale Structured State Space Models
Mohammad Shahab Sepehri, Zalan Fabian, and Mahdi Soltanolkotabi. 2024 · 2024
Closest in time.
Integrating mamba sequence model and hierarchical upsampling network for accurate semantic segmentation of multiple sclerosis lesion. In International Conference on Life System Modeling and Simulation . Springer, 365–379
Kazi Shahriar Sanjid, Md Tanzim Hossain, Md Shakib Shahariar Junayed, Mohammad Monir Uddin, Yu-Long Wang, and Nasir M Uddin. 2024 · 2024
Closest in time.
SSAMBA: Self-Supervised Audio Representation Learning With Mamba State Space Model. In 2024 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 1053–1059
Siavash Shams, Sukru Samet Dindar, Xilin Jiang, and Nima Mesgarani. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Md Atik Ahamed and Qiang Cheng. 2024 · 2024
Cited alongside, same era.
BlackMamba: Mixture of Experts for State-Space Models. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models
Quentin Gregory Anthony, Yury Tokpanov, Paolo Glorioso, and Beren Millidge. 2024 · 2024
Cited alongside, same era.
Graph mamba: Towards learning on graphs with state space models. In Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining . 119–130
Ali Behrouz and Farnoosh Hashemi. 2024 · 2024
Cited alongside, same era.
Mambamixer: Efficient selective state space models with dual token and channel selection
Ali Behrouz, Michele Santacatterina, and Ramin Zabih. 2024 · 2024
Cited alongside, same era.
DASS: Distilled audio state space models are stronger and more duration-scalable learners. In 2024 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 1015–1022
Saurabhchand Bhati, Yuan Gong, Leonid Karlinsky, Hilde Kuehne, Rogerio Feris, and James Glass. 2024 · 2024
Cited alongside, same era.
Hierarchical State Space Models for Continuous Sequence-to-Sequence Modeling. In International Conference on Machine Learning . PMLR, 3795–3816
Raunaq Bhirangi, Chenyu Wang, Venkatesh Pattabiraman, Carmel Majidi, Abhinav Gupta, Tess Hellebrekers, and Lerrel Pinto. 2024 · 2024
Cited alongside, same era.
A Mamba-Based Foundation Model for Chemistry. In Neurips 2024 Workshop Foundation Models for Science: Progress, Opportunities, and Challenges
Emilio Vital Brazil, Eduardo Soares, Victor Yukio Shirasuna, Renato Cerqueira, Dmitry Zubarev, and Kristin Schmidt. [n. d.] · 2024
Cited alongside, same era.
Mamba as Decision Maker: Exploring Multi-scale Sequence Modeling in Offline Reinforcement Learning
Jiahang Cao, Qiang Zhang, Ziqing Wang, Jiaxu Wang, Hao Cheng, Yecheng Shao, Wen Zhao, Gang Han, Yijie Guo, and Renjing Xu. 2024 · 2024
Cited alongside, same era.
DualMamba: A lightweight spectral–spatial mamba-convolution network for hyperspectral image classification
Jiamu Sheng, Jingyi Zhou, Jiong Wang, Peng Ye, and Jiayuan Fan. 2024 · 2024
Closest in time.
Multi-scale vmamba: Hierarchy in hierarchy visual state space model
Yuheng Shi, Minjing Dong, and Chang Xu. 2024 · 2024
Closest in time.
Freeu: Free lunch in diffusion u-net. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4733–4743
Chenyang Si, Ziqi Huang, Yuming Jiang, and Ziwei Liu. 2024 · 2024
Closest in time.
Tramba: A hybrid transformer and mamba architecture for practical audio and bone conduction speech super resolution and enhancement on mobile and wearable platforms
Yueyuan Sui, Minghui Zhao, Junxi Xia, Xiaofan Jiang, and Stephen Xia. 2024 · 2024
Closest in time.
Vmrnn: Integrating vision mamba and lstm for efficient and accurate spatiotemporal forecasting. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5663–5673
Yujin Tang, Peijie Dong, Zhenheng Tang, Xiaowen Chu, and Junwei Liang. 2024 · 2024
Closest in time.
MSAMamba: Adapting Subquadratic Sequence Models to Long-Context DNA MSA Analysis. In The 14th International Conference on Information Science and Technology . 1–10
Vishrut Thoutam and Dina Ellsworth. 2024 · 2024
Closest in time.
An Empirical Study of Mamba-based Language Models
Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, Sudhakar Singh, Deepak Narayanan, et al · 2024
Closest in time.
Graph-mamba: Towards long-range graph sequence modeling with selective state spaces
Chloe Wang, Oleksii Tsepa, Jun Ma, and Bo Wang. 2024e · 2024
Closest in time.
Memorymamba: Memory-augmented state space model for defect recognition
Qianning Wang, He Hu, and Yucheng Zhou. 2024d · 2024
Closest in time.
State space model for new-generation network alternative to transformers: A survey
Xiao Wang, Shiao Wang, Yuhe Ding, Yuehang Li, Wentao Wu, Yao Rong, Weizhe Kong, Ju Huang, Shihao Li, Haoxiang Yang, et al · 2024
Closest in time.
Yuda Wang, Xuxin He, and Shengxin Zhu. 2024c · 2024
Closest in time.
Ziyang Wang, Jian-Qing Zheng, Chao Ma, and Tao Guo. 2024g · 2024
Closest in time.
ProMamba: Prompt-Mamba for polyp segmentation
Jianhao Xie, Ruofan Liao, Ziang Zhang, Sida Yi, Yuesheng Zhu, and Guibo Luo. 2024b · 2024
Closest in time.
Fusionmamba: Dynamic feature enhancement for multimodal image fusion with mamba
Xinyu Xie, Yawen Cui, Chio-In Ieong, Tao Tan, Xiaozhi Zhang, Xubin Zheng, and Zitong Yu. 2024a · 2024
Closest in time.
Smiles-mamba: Chemical mamba foundation models for drug admet prediction
Bohao Xu, Yingzhou Lu, Chenhao Li, Ling Yue, Xiao Wang, Nan Hao, Tianfan Fu, and Jim Chen. 2024a · 2024
Closest in time.
Visual mamba: A survey and new outlooks
Rui Xu, Shu Yang, Yihui Wang, Yu Cai, Bo Du, and Hao Chen. 2024b · 2024
Closest in time.
RankMamba, Benchmarking Mamba’s Document Ranking Performance in the Era of Transformers
Zhichao Xu. 2024 · 2024
Closest in time.
Guangqian Yang, Kangrui Du, Zhihan Yang, Ye Du, Yongping Zheng, and Shujun Wang. 2024b · 2024
Closest in time.
Uncovering Selective State Space Model’s Capabilities in Lifelong Sequential Recommendation
Jiyuan Yang, Yuanzi Li, Jingyu Zhao, Hanbing Wang, Muyang Ma, Jun Ma, Zhaochun Ren, Mengqi Zhang, Xin Xin, Zhumin Chen, et al · 2024
Closest in time.
Judy X Yang, Jun Zhou, Jing Wang, Hui Tian, and Alan Wee Chung Liew. 2024e · 2024
Closest in time.
Vivim: A video vision mamba for medical video segmentation
Yijun Yang, Zhaohu Xing, Lequan Yu, Chunwang Huang, Huazhu Fu, and Lei Zhu. 2024d · 2024
Closest in time.
Spectralmamba: Efficient mamba for hyperspectral image classification
Jing Yao, Danfeng Hong, Chenyu Li, and Jocelyn Chanussot. 2024 · 2024
Closest in time.
Zi Ye and Tianxiang Chen. 2024 · 2024
Closest in time.
Mvgamba: Unify 3d content generation as state space sequence modeling
Xuanyu Yi, Zike Wu, Qiuhong Shen, Qingshan Xu, Pan Zhou, Joo-Hwee Lim, Shuicheng Yan, Xinchao Wang, and Hanwang Zhang. 2024 · 2024
Closest in time.
Medmamba: Vision mamba for medical image classification
Yubiao Yue and Zhenzhang Li. 2024 · 2024
Closest in time.
Mambamos: Lidar-based 3d moving object segmentation with motion-aware state space model. In Proceedings of the 32nd ACM International Conference on Multimedia . 1505–1513
Kang Zeng, Hao Shi, Jiacheng Lin, Siyu Li, Jintao Cheng, Kaiwei Wang, Zhiyong Li, and Kailun Yang. 2024 · 2024
Closest in time.
Deep learning models for price forecasting of financial time series: A review of recent advancements: 2020–2022
Cheng Zhang, Nilam Nur Amir Sjarif, and Roslina Ibrahim. 2024c · 2024
Closest in time.
A survey on visual mamba
Hanwei Zhang, Ying Zhu, Dan Wang, Lijun Zhang, Tianxiang Chen, Ziyang Wang, and Zi Ye. 2024f · 2024
Closest in time.
Linear-Time Graph Neural Networks for Scalable Recommendations. In Proceedings of the ACM on Web Conference 2024 . 3533–3544
Jiahao Zhang, Rui Xue, Wenqi Fan, Xin Xu, Qing Li, Jian Pei, and Xiaorui Liu. 2024d · 2024
Closest in time.
MAMC—Optimal on accuracy and efficiency for automatic modulation classification with extended signal length
Yezhuo Zhang, Zinan Zhou, Yichao Cao, Guangyu Li, and Xuanpeng Li. 2024e · 2024
Closest in time.
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
Zeyu Zhang, Akide Liu, Qi Chen, Feng Chen, Ian Reid, Richard Hartley, Bohan Zhuang, and Hao Tang. 2024a · 2024
Closest in time.
Rs-mamba for large remote sensing image dense prediction
Sijie Zhao, Hao Chen, Xueliang Zhang, Pengfeng Xiao, Lei Bai, and Wanli Ouyang. 2024a · 2024
Closest in time.
Recommender systems in the era of large language models (llms)
Zihuai Zhao, Wenqi Fan, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Zhen Wen, Fei Wang, Xiangyu Zhao, Jiliang Tang, et al · 2024
Closest in time.
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model. In International Conference on Machine Learning . PMLR, 62429–62442
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024 · 2024
Closest in time.
TSCMamba: Mamba meets multi-view learning for time series classification
Md Atik Ahamed and Qiang Cheng. 2025 · 2025
Closest in time.
Damba-ST: Domain-Adaptive Mamba for Efficient Urban Spatio-Temporal Prediction
Rui An, Yifeng Zhang, Ziran Liang, Wenqi Fan, Yuxuan Liang, Xuequn Shang, and Qing Li. 2025 · 2025
Closest in time.
P-SPIKESSM: HARNESSING PROBABILISTIC SPIKING STATE SPACE MODELS FOR LONG-RANGE DEPENDENCY TASKS. In 13th International Conference on Learning Representations, ICLR 2025 . 77267–77282
Malyaban Bal and Abhronil Sengupta. 2025 · 2025
Closest in time.
How Do Large Language Models Understand Graph Patterns? A Benchmark for Graph Pattern Comprehension. In The Thirteenth International Conference on Learning Representations . 25917–25947
Xinnan Dai, Haohao Qu, Yifei Shen, Bohang Zhang, Qihao Wen, Wenqi Fan, Dongsheng Li, Jiliang Tang, and Caihua Shan. 2025 · 2025
Closest in time.
Fusion-Mamba for Cross-Modality Object Detection
Wenhao Dong, Haodong Zhu, Shaohui Lin, Xiaoyan Luo, Yunhang Shen, Guodong Guo, and Baochang Zhang. 2025 · 2025
Closest in time.
nnmamba: 3d biomedical image segmentation, classification and landmark detection with state space model. In 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI) . IEEE, 1–5
Haifan Gong, Luoyao Kang, Yitao Wang, Yihan Wang, Xiang Wan, Xusheng Wu, and Haofeng Li. 2025 · 2025
Closest in time.
Demystify Mamba in Vision: A Linear Attention Perspective
Dongchen Han, Ziyi Wang, Zhuofan Xia, Yizeng Han, Yifan Pu, Chunjiang Ge, Jun Song, Shiji Song, Bo Zheng, and Gao Huang. 2025 · 2025
Closest in time.
Mambavision: A hybrid mamba-transformer vision backbone. In Proceedings of the Computer Vision and Pattern Recognition Conference . 25261–25270
Ali Hatamizadeh and Jan Kautz. 2025 · 2025
Closest in time.
DenseSSM: State Space Models with Dense Hidden Connection for Efficient Large Language Models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) . 9243–9254
Wei He, Kai Han, Yehui Tang, Chengcheng Wang, Yujie Yang, Tianyu Guo, and Yunhe Wang. 2025b · 2025
Closest in time.
Pan-mamba: Effective pan-sharpening with state space model
Xuanhua He, Ke Cao, Jie Zhang, Keyu Yan, Yingying Wang, Rui Li, Chengjun Xie, Danfeng Hong, and Man Zhou. 2025a · 2025
Closest in time.
Sum: Saliency unification through mamba for visual attention modeling. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) . IEEE, 1597–1607
Alireza Hosseini, Amirhossein Kazerouni, Saeed Akhavan, Michael Brudno, and Babak Taati. 2025 · 2025
Closest in time.
Self-prior guided Mamba-UNet networks for medical image super-resolution. In International Conference on Pattern Recognition . Springer, 160–174
Zexin Ji, Beiji Zou, Xiaoyan Kui, Pierre Vera, and Su Ruan. 2025 · 2025
Closest in time.
Dual-path mamba: Short and long-term bidirectional selective structured state space models for speech separation. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
Xilin Jiang, Cong Han, and Nima Mesgarani. 2025 · 2025
Closest in time.
Jamba: Hybrid transformer-mamba language models. In The thirteenth international conference on learning representations . 48624–48649
Barak Lenz, Opher Lieber, Alan Arazi, Amir Bergman, Avshalom Manevich, Barak Peleg, Ben Aviram, Chen Almagor, Clara Fridman, Dan Padnos, et al · 2025
Closest in time.
Vision mamba: A comprehensive survey and taxonomy
Xiao Liu, Chenxu Zhang, Fuxiang Huang, Shuyin Xia, Guoyin Wang, and Lei Zhang. 2025 · 2025
Closest in time.
Towards evaluating the robustness of visual state space models. In Proceedings of the Computer Vision and Pattern Recognition Conference . 3544–3553
Hashmat Shadab Malik, Fahad Shamshad, Muzammal Naseer, Karthik Nandakumar, Fahad Shahbaz Khan, and Salman Khan. 2025 · 2025
Closest in time.
A survey of webagents: Towards next-generation ai agents for web automation with large foundation models. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2 . 6140–6150
Liangbo Ning, Ziran Liang, Zhuohang Jiang, Haohao Qu, Yujuan Ding, Wenqi Fan, Xiao-yong Wei, Shanru Lin, Hui Liu, Philip S Yu, et al · 2025
Closest in time.
Mamba-360: Survey of state space models as transformer alternative for long sequence modelling: Methods, applications, and challenges
Badri Narayana Patro and Vijay Srinivas Agneeswaran. 2025 · 2025
Closest in time.
Efficientvmamba: Atrous selective scan for light weight visual mamba. In Proceedings of the AAAI conference on artificial intelligence , Vol. 39. 6443–6451
Xiaohuan Pei, Tao Huang, and Chang Xu. 2025 · 2025
Closest in time.
TokenRec: Learning to Tokenize ID for LLM-Based Generative Recommendations
Haohao Qu, Wenqi Fan, Zihuai Zhao, and Qing Li. 2025a · 2025
Closest in time.
Diffusion Generative Recommendation with Continuous Tokens
Haohao Qu, Shanru Lin, Yujuan Ding, Yiqi Wang, and Wenqi Fan. 2025b · 2025
Closest in time.
ProtMamba: a homology-aware but alignment-free protein state space model
Damiano Sgarbossa, Cyril Malbranke, and Anne-Florence Bitbol. 2025 · 2025
Closest in time.
Vmambair: Visual state space model for image restoration
Yuan Shi, Bin Xia, Xiaoyu Jin, Xing Wang, Tianyu Zhao, Xin Xia, Xuefeng Xiao, and Wenming Yang. 2025 · 2025
Closest in time.
Mlsa4rec: Mamba combined with low-rank decomposed self-attention for sequential recommendation
Jinzhao Su, Zhenhua Huang, Chang-Dong Wang, and Yunwen Chen. 2025 · 2025
Closest in time.
Sigma: Siamese mamba network for multi-modal semantic segmentation. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) . IEEE, 1734–1744
Zifu Wan, Pingping Zhang, Yuhao Wang, Silong Yong, Simon Stepputtis, Katia Sycara, and Yaqi Xie. 2025 · 2025
Closest in time.
Graph machine learning in the era of large language models (llms)
Shijie Wang, Jiani Huang, Zhikai Chen, Yu Song, Wenzhuo Tang, Haitao Mao, Wenqi Fan, Hui Liu, Xiaorui Liu, Dawei Yin, et al · 2025
Closest in time.
Text-controlled motion mamba: Text-instructed temporal grounding of human motion
Xinghan Wang, Zixi Kang, and Yadong Mu. 2025b · 2025
Closest in time.
Ultralight vm-unet: Parallel vision mamba significantly reduces parameters for skin lesion segmentation
Renkai Wu, Yinghao Liu, Guochen Ning, Pengchen Liang, and Qing Chang. 2025 · 2025
Closest in time.
Segmamba-v2: Long-range sequential modeling mamba for general 3d medical image segmentation
Zhaohu Xing, Tian Ye, Yijun Yang, Du Cai, Baowen Gai, Xiao-Jian Wu, Feng Gao, and Lei Zhu. 2025 · 2025
Closest in time.
SST: Multi-Scale Hybrid Mamba-Transformer Experts for Time Series Forecasting. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management . 3655–3665
Xiongxiao Xu, Canyu Chen, Yueqing Liang, Baixiang Huang, Guangji Bai, Liang Zhao, and Kai Shu. 2025 · 2025
Closest in time.
SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering
Zhe Yang, Wenrui Li, and Guanghui Cheng. 2025 · 2025
Closest in time.
SSD4Rec: A Structured State Space Duality Model for Efficient Sequential Recommendation
Yifeng Zhang, Haohao Qu, Liangbo Ning, Wenqi Fan, and Qing Li. 2025a · 2025
Closest in time.
Cobra: Extending mamba to multi-modal large language model for efficient inference. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 10421–10429
Han Zhao, Min Zhang, Wei Zhao, Pengxiang Ding, Siteng Huang, and Donglin Wang. 2025 · 2025
Closest in time.
Yi Zhou, Haohao Qu, Yunqing Liu, Shanru Lin, Le Song, and Wenqi Fan. 2025a · 2025
Closest in time.
Rhythmmamba: Fast, lightweight, and accurate remote physiological measurement. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 11077–11085
Bochao Zou, Zizheng Guo, Xiaocheng Hu, and Huimin Ma. 2025 · 2025
Closest in time.
DeMa: Dual-Path Delay-Aware Mamba for Efficient Multivariate Time Series Analysis
Rui An, Haohao Qu, Wenqi Fan, Xuequn Shang, and Qing Li. 2026 · 2026
Closest in time.
Weak-Mamba-UNet: Visual Mamba Makes CNN and ViT Work Better for Scribble-based Medical Image Segmentation
Ziyang Wang, Tianli Tao, Yiyuan Ge, Zhihao Chen, Tianxiang Chen, Tianxiang Chen, Zi Ye, and Yongxiang Lei. 2026 · 2026
Closest in time.
A graph neural network framework for social recommendations
Wenqi Fan, Yao Ma, Qing Li, Jianping Wang, Guoyong Cai, Jiliang Tang, and Dawei Yin. 2020 · 2047
Closest in time.