Fetching the paper…
Reading the bibliography…
State-space models (SSMs), particularly the Mamba architecture, have emerged as powerful alternatives to Transformers for sequence modeling, offering linear-time complexity and competitive performance across diverse tasks.
Language models are few-shot learners
Tom Brown and 1 others. 2020 · 1901
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla. 1990 · 1990
Earlier work this paper cites.
Optimal brain surgeon: Extensions and performance comparisons
Babak Hassibi, David G Stork, and Gregory J Wolff. 1993 · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Song Han, Jeff Pool, John Tran, and William J Dally. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J Dally. 2016 · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. 2016 · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Michael H Zhu and Suyog Gupta. 2017 · 2017
Earlier work this paper cites.
Deep rewiring: Training very sparse deep networks
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert Legenstein. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Earlier work this paper cites.
Amc: Automl for model compression and acceleration on mobile devices
Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han. 2018 · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018 · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr. 2018 · 2018
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich. 2019 · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
Importance estimation for neural network pruning
Pavlo Molchanov, Arun Mallya, Stephen Tyree, Irene Frosio, and Jan Kautz. 2019 · 2019
Cited alongside, same era.
Energy and policy considerations for deep learning in nlp
Practical issues in recurrent neural network pruning
Shaobo Liu, Ang Li, Bowen Feng, Feng Zhang, and Zewen Zhang. 2021 · 2021
Later among the works it cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021 · 2021
Later among the works it cites.
S4nd: Modeling images and videos as multidimensional signals using state spaces
Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré. 2022 · 2022
Later among the works it cites.
Diagonal state spaces are as effective as structured state spaces
Ankit Gupta, Aditya Kusupati, Haricharan Simhadri, Harsh Pathak, Aniruddha Kembhavi, and Jitendra Malik. 2022 · 2022
Later among the works it cites.
Liquid structural state-space models
Ramin Hasani, Mathias Lechner, Alexander Amini, Lucas Liebenwein, Max Tschernuth, Joshua Tenenbaum, and Daniela Rus. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Model compression and hardware acceleration for neural networks: A comprehensive survey
Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie. 2020 · 2020
Cited alongside, same era.
Pretrained transformers improve out-of-distribution robustness
Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. 2020 · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander M Rush. 2020 · 2020
Cited alongside, same era.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
S4nd: Modeling images and videos as multidimensional signals using state spaces
Tri Nguyen, Smit Baguley, Anthony Dao, Karan Goel, Hamza Hassani, and Albert Gu. 2022 · 2022
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. 2023 · 2023
Later among the works it cites.
Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution
Jason Nguyen, Gunnar Rätsch Azis, and Konstantin Rohr. 2023 · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Kamalraj Kanakarajan, and Mathieu Ravaut. 2023 · 2023
Later among the works it cites.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Hugo Larochelle, and Stefano Ermon. 2023 · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Yutao Sun, Li Dong, Xipeng Qiu, and Tianyu Yang. 2023 · 2023
Later among the works it cites.
Transformers and state space models: A detailed analysis through the lens of spectral properties
Tri Dao, Shiyi Fu, Zhengxin Chen, Orhan Firat, Kwonjoon Lee, and Albert Gu. 2024 · 2024
Later among the works it cites.
Luna: Language understanding with natural state space models
Weizhe Ma, Sharan Narang, Sreyas Bodapati, Anirudh Madaan, Hao Zhou, Daniel Kirchner, Christopher Akiki, Helen Firoozi, Kathy Lee, Collin Schulman, and 1 others. 2024 · 2024
Later among the works it cites.
Vl-mamba: Exploring state space models for multimodal learning
Yanyuan Qiao, Donghai Gong, Aozhu Liu, Chunyuan Li, Yin Bi, Ting Yao, Wei Chen, and Dongmei Zhang. 2024 · 2024
Later among the works it cites.
Simple and effective gradient-based pruning for large language models
Jiawei Sun, Husheng Wang, Shuai Liu, Ke Ren, Mingyu Gao, Chenjuan Xu, and Bin Guo. 2024 · 2024
Later among the works it cites.
Mambaformer: Efficient language modeling with selective state spaces and attention
Szymon Tworkowski, Konrad Przybysz, Tomasz Korbak, Pedro Rodriguez, Wojciech Rozemłyn, Kamil Rakowski, Maciej Grzelak, Albert Webson, Maciej Szafraniec, Robin Sorsch, and 1 others. 2024 · 2024
Later among the works it cites.
Vision mamba: Efficient visual representation learning with bidirectional state space model
Yupeng Zhu, Huang Hu, Dongchen He, Zhe Gan, Zhiqi Wang, Lijuan Wang, Zicheng Zhang, Jianfeng Liu, Guangxing Cheng, Qi Tian, Dacheng Yu, and Xinggang Wang. 2024 · 2024
Later among the works it cites.