Fetching the paper…
Reading the bibliography…
Transformer, a deep neural network architecture, has long dominated the field of natural language processing and beyond.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman · 1960
Earlier work this paper cites.
A model for reasoning about persistence and causation
Thomas Dean and Keiji Kanazawa · 1989
Earlier work this paper cites.
Modern control theory
William L Brogan · 1991
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Time series analysis by state space methods , volume 38
James Durbin and Siem Jan Koopman · 2012
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Transformer dissection: a unified understanding of transformer’s attention via the lens of kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Graph transformer networks
Seongjun Yun, Minbyul Jeong, Raehyun Kim, Jaewoo Kang, and Hyunwoo J Kim · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Message passing neural networks
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl · 2020
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
Earlier work this paper cites.
Framing rnn as a kernel method: A neural ode approach
Adeline Fermanian, Pierre Marion, Jean-Philippe Vert, and Gérard Biau · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
Long range graph benchmark
Vijay Prakash Dwivedi, Ladislav Rampášek, Michael Galkin, Ali Parviz, Guy Wolf, Anh Tuan Luu, and Dominique Beaini · 2022
Earlier work this paper cites.
Hungry hungry hippos: Towards language modeling with state space models
Daniel Y Fu, Tri Dao, Khaled K Saab, Armin W Thomas, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
How to train your hippo: State space models with generalized orthogonal basis projections
Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
Recipe for a general, powerful, scalable graph transformer
Ladislav Rampášek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Dominique Beaini · 2022
Earlier work this paper cites.
Junxiong Wang, Jing Nathan Yan, Albert Gu, and Alexander M Rush · 2022
Earlier work this paper cites.
Efficient long sequence modeling via state space augmented transformer
Simiao Zuo, Xiaodong Liu, Jian Jiao, Denis Charles, Eren Manavoglu, Tuo Zhao, and Jianfeng Gao · 2022
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Earlier work this paper cites.
On over-squashing in message passing neural networks: The impact of width, depth, and topology
Francesco Di Giovanni, Lorenzo Giusti, Federico Barbero, Giulia Luise, Pietro Lio, and Michael M Bronstein · 2023
Earlier work this paper cites.
Mahan Fathi, Jonathan Pilault, Pierre-Luc Bacon, Christopher Pal, Orhan Firat, and Ross Goroshin · 2023
Earlier work this paper cites.
Simple hardware-efficient long convolutions for sequence modeling
Daniel Y Fu, Elliot L Epstein, Eric Nguyen, Armin W Thomas, Michael Zhang, Tri Dao, Atri Rudra, and Christopher Ré · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Earlier work this paper cites.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Earlier work this paper cites.
A survey on oversmoothing in graph neural networks
T Konstantin Rusch, Michael M Bronstein, and Siddhartha Mishra · 2023
Earlier work this paper cites.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou · 2023
Earlier work this paper cites.
Blackmamba: Mixture of experts for state-space models
Quentin Anthony, Yury Tokpanov, Paolo Glorioso, and Beren Millidge · 2024
Earlier work this paper cites.
Vim-unet: Vision mamba for biomedical segmentation, 2024
Anwai Archit and Constantin Pape · 2024
Cited alongside, same era.
Simple linear attention language models balance the recall-throughput tradeoff
Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré · 2024
Cited alongside, same era.
Retinexmamba: Retinex-based mamba for low-light image enhancement, 2024
Jiesong Bai, Yuhao Yin, Qiyuan He, Yuanxian Li, and Xiaofeng Zhang · 2024
Cited alongside, same era.
Graph mamba: Towards learning on graphs with state space models
Ali Behrouz and Farnoosh Hashemi · 2024
Cited alongside, same era.
Locost: State-space models for long document abstractive summarization
Florian Le Bronnec, Song Duong, Mathieu Ravaut, Alexandre Allauzen, Nancy F Chen, Vincent Guigue, Alberto Lumbreras, Laure Soulier, and Patrick Gallinari · 2024
Efficientvmamba: Atrous selective scan for light weight visual mamba
Xiaohuan Pei, Tao Huang, and Chang Xu · 2024
Closest in time.
Fusionmamba: Efficient image fusion with state space model
Siran Peng, Xiangyu Zhu, Haoyu Deng, Zhen Lei, and Liang-Jian Deng · 2024
Closest in time.
Moe-mamba: Efficient selective state space models with mixture of experts
Maciej Pióro, Kamil Ciebiera, Krystian Król, Jan Ludziejewski, and Sebastian Jaszczur · 2024
Closest in time.
Smcd: High realism motion style transfer via mamba-based diffusion, 2024
Ziyun Qian, Zeyu Xiao, Zhenyi Wu, Dingkang Yang, Mingcheng Li, Shunli Wang, Shuaibing Wang, Dongliang Kou, and Lihua Zhang · 2024
Closest in time.
Vl-mamba: Exploring state space models for multimodal learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A novel state space model with local enhancement and state sharing for image fusion
Zihan Cao, Xiao Wu, Liang-Jian Deng, and Yu Zhong · 2024
Cited alongside, same era.
An investigation of incorporating mamba for speech enhancement, 2024
Rong Chao, Wen-Huang Cheng, Moreno La Quatra, Sabato Marco Siniscalchi, Chao-Han Huck Yang, Szu-Wei Fu, and Yu Tsao · 2024
Cited alongside, same era.
Simba: Mamba augmented u-shiftgcn for skeletal action recognition in videos, 2024
Soumyabrata Chaudhuri and Saumik Bhattacharya · 2024
Cited alongside, same era.
Activating wider areas in image super-resolution
Cheng Cheng, Hang Wang, and Hongbin Sun · 2024
Cited alongside, same era.
Feasibility of state space models for network traffic generation, 2024
Andrew Chu, Xi Jiang, Shinan Liu, Arjun Bhagoji, Francesco Bronzino, Paul Schmitt, and Nick Feamster · 2024
Cited alongside, same era.
Tri Dao and Albert Gu · 2024
Cited alongside, same era.
Griffin: Mixing gated linear recurrences with local attention for efficient language models
Soham De, Samuel L Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, et al · 2024
Cited alongside, same era.
Yanyuan Qiao, Zheng Yu, Longteng Guo, Sihan Chen, Zijia Zhao, Mingzhen Sun, Qi Wu, and Jing Liu · 2024
Closest in time.
Multichannel long-term streaming neural speech enhancement for static and moving speakers
Changsheng Quan and Xiaofei Li · 2024
Closest in time.
Samba: Simple hybrid state space models for efficient unlimited context language modeling, 2024
Liliang Ren, Yang Liu, Yadong Lu, Yelong Shen, Chen Liang, and Weizhu Chen · 2024
Closest in time.
Vm-unet: Vision mamba unet for medical image segmentation
Jiacheng Ruan and Suncheng Xiang · 2024
Closest in time.
Caduceus: Bi-directional equivariant long-range dna sequence modeling
Yair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao, Albert Gu, and Volodymyr Kuleshov · 2024
Closest in time.
Ssamba: Self-supervised audio representation learning with mamba state space model, 2024
Siavash Shams, Sukru Samet Dindar, Xilin Jiang, and Nima Mesgarani · 2024
Closest in time.
Gamba: Marry gaussian splatting with mamba for single view 3d reconstruction, 2024
Qiuhong Shen, Zike Wu, Xuanyu Yi, Pan Zhou, Hanwang Zhang, Shuicheng Yan, and Xinchao Wang · 2024
Closest in time.
Dualmamba: A lightweight spectral-spatial mamba-convolution network for hyperspectral image classification, 2024
Jiamu Sheng, Jingyi Zhou, Jiong Wang, Peng Ye, and Jiayuan Fan · 2024
Closest in time.
Mambastock: Selective state space model for stock prediction
Zhuangwei Shi · 2024
Closest in time.
Rotate to scan: Unet-like mamba with triplet ssm module for medical image segmentation, 2024
Hao Tang, Lianglun Cheng, Guoheng Huang, Zhengguang Tan, Junhao Lu, and Kaihong Wu · 2024
Closest in time.
Dim: Diffusion mamba for efficient high-resolution image synthesis
Yao Teng, Yue Wu, Han Shi, Xuefei Ning, Guohao Dai, Yu Wang, Zhenguo Li, and Xihui Liu · 2024
Closest in time.
Uu-mamba: Uncertainty-aware u-mamba for cardiac image segmentation, 2024
Ting Yu Tsai, Li Lin, Shu Hu, Ming-Ching Chang, Hongtu Zhu, and Xin Wang · 2024
Closest in time.
Soar: Advancements in small body object detection for aerial imagery using state space models and programmable gradients, 2024
Tushar Verma, Jyotsna Singh, Yash Bhartari, Rishi Jarwal, Suraj Singh, and Shubhkarman Singh · 2024
Closest in time.
Sigma: Siamese mamba network for multi-modal semantic segmentation, 2024
Zifu Wan, Yuhao Wang, Silong Yong, Pingping Zhang, Simon Stepputtis, Katia Sycara, and Yaqi Xie · 2024
Closest in time.
Pointabm:integrating bidirectional state space model with multi-head self-attention for point cloud analysis, 2024
Jia wei Chen, Yu jie Xiong, and Yong bin Gao · 2024
Closest in time.
H-vmunet: High-order vision mamba unet for medical image segmentation
Renkai Wu, Yinghao Liu, Pengchen Liang, and Qing Chang · 2024
Closest in time.
Lfmamba: Light field image super-resolution with state space model, 2024
Wang xia, Yao Lu, Shunzhou Wang, Ziqi Wang, Peiqi Xia, and Tianfei Zhou · 2024
Closest in time.
Promamba: Prompt-mamba for polyp segmentation, 2024
Jianhao Xie, Ruofan Liao, Ziang Zhang, Sida Yi, Yuesheng Zhu, and Guibo Luo · 2024
Closest in time.
Segmamba: Long-range sequential modeling mamba for 3d medical image segmentation
Zhaohu Xing, Tian Ye, Yijun Yang, Guang Liu, and Lei Zhu · 2024
Closest in time.
Spectralmamba: Efficient mamba for hyperspectral image classification, 2024
Jing Yao, Danfeng Hong, Chenyu Li, and Jocelyn Chanussot · 2024
Closest in time.
Mambaout: Do we really need mamba for vision?
Weihao Yu and Xinchao Wang · 2024
Closest in time.
Mucm-net: A mamba powered ucm-net for skin lesion segmentation, 2024
Chunyu Yuan, Dongfang Zhao, and Sos S. Agaian · 2024
Closest in time.
Medmamba: Vision mamba for medical image classification
Yubiao Yue and Zhenzhang Li · 2024
Closest in time.
C-mamba: Channel correlation enhanced state space models for multivariate time series forecasting, 2024
Chaolv Zeng, Zhanyu Liu, Guanjie Zheng, and Linghe Kong · 2024
Closest in time.
Cobra: Extending mamba to multi-modal large language model for efficient inference
Han Zhao, Min Zhang, Wei Zhao, Pengxiang Ding, Siteng Huang, and Donglin Wang · 2024
Closest in time.
Freqmamba: Viewing mamba from a frequency perspective for image deraining, 2024
Zou Zhen, Yu Hu, and Zhao Feng · 2024
Closest in time.
U-shaped vision mamba for single image dehazing
Zhuoran Zheng and Chen Wu · 2024
Closest in time.
Fd-vision mamba for endoscopic exposure correction
Zhuoran Zheng and Jun Zhang · 2024
Closest in time.