Fetching the paper…
Reading the bibliography…
While Transformers have been the main architecture behind deep learning's success in language modeling, state-space models (SSMs) such as Mamba have recently been shown to match or outperform Transformers at small to medium scale.
“Language Models are Few-shot Learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“Data Parallel Algorithms”
W Hillis and Guy Steele · 1986
Earlier work this paper cites.
“Prefix Sums and Their Applications”
Guy Blelloch · 1990
Earlier work this paper cites.
“Pade Approximants: Encyclopedia of Mathematics and It’s Applications, Vol. 59 George A. Baker, Jr., Peter Graves-Morris”
George Baker, George Baker, Peter Graves-Morris and Susan Baker · 1996
Earlier work this paper cites.
“Long Short-Term Memory”
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“On a new class of structured matrices”
Yuli Eidelman and Israel Gohberg · 1999
Earlier work this paper cites.
“A bibliography on semiseparable matrices”
Raf Vandebril, M Barel, Gene Golub and Nicola Mastronardi · 2005
Earlier work this paper cites.
“Random Features for Large-Scale Kernel Machines”
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
“Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling”
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho and Yoshua Bengio · 2014
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate”
Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio · 2015
Earlier work this paper cites.
“Time Series Analysis: Forecasting and Control”
George Box, Gwilym Jenkins, Gregory Reinsel and Greta Ljung · 2015
Earlier work this paper cites.
“Quasi-recurrent Neural Networks”
James Bradbury, Stephen Merity, Caiming Xiong and Richard Socher · 2016
Earlier work this paper cites.
“Gaussian Error Linear Units (GELUs)”
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
“The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context”
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda and Raquel Fernández · 2016
Earlier work this paper cites.
“Computing with Quasiseparable Matrices”
Clément Pernet · 2016
Earlier work this paper cites.
“Simple Recurrent Units for Highly Parallelizable Recurrence”
Tao Lei, Yu Zhang, Sida Wang, Hui Dai and Yoav Artzi · 2017
Earlier work this paper cites.
“Swish: A Self-gated Activation Function”
Prajit Ramachandran, Barret Zoph and Quoc Le · 2017
Earlier work this paper cites.
“Attention Is All You Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan. Gomez, Lukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick and Oyvind Tafjord · 2018
Earlier work this paper cites.
“A Two-Pronged Progress in Structured Dense Matrix Vector Multiplication”
Christopher De, Albert Gu, Rohan Puttagunta, Christopher Ré and Atri Rudra · 2018
Earlier work this paper cites.
“Parallelizing Linear Recurrent Neural Nets Over Sequence Length”
Eric Martin and Chris Cundy · 2018
Earlier work this paper cites.
“Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering”
Todor Mihaylov, Peter Clark, Tushar Khot and Ashish Sabharwal · 2018
Earlier work this paper cites.
“Time and space efficient generators for quasiseparable matrices”
Clément Pernet and Arne Storjohann · 2018
Earlier work this paper cites.
“Learning Compressed Transforms with Low Displacement Rank”
Anna Thomas, Albert Gu, Tri Dao, Atri Rudra and Christopher Ré · 2018
Earlier work this paper cites.
“Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations”
Tri Dao, Albert Gu, Matthew Eichhorn, Atri Rudra and Christopher Ré · 2019
Earlier work this paper cites.
“Fast Transformer Decoding: One Write-head is All You Need”
Noam Shazeer · 2019
Earlier work this paper cites.
“Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism”
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper and Bryan Catanzaro · 2019
Earlier work this paper cites.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi and Yejin Choi · 2019
Earlier work this paper cites.
“PIQA: Reasoning about Physical Commonsense in Natural Language”
Yonatan Bisk, Rowan Zellers, Jianfeng Gao and Yejin Choi · 2020
Earlier work this paper cites.
“Kaleidoscope: An Efficient, Learnable Representation for All Structured Linear Maps”
Tri Dao, Nimit Sohoni, Albert Gu, Matthew Eichhorn, Amit Blonder, Megan Leszczynski, Atri Rudra and Christopher Ré · 2020
Earlier work this paper cites.
“The Pile: An 800GB Dataset of Diverse Text for Language Modeling”
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser and Connor Leahy · 2020
Earlier work this paper cites.
“HIPPO: Recurrent Memory with Optimal Polynomial Projections”
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra and Christopher Ré · 2020
Earlier work this paper cites.
“Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and François Fleuret · 2020
Earlier work this paper cites.
“Linear Dynamical Systems as a Core Computational Primitive”
Shiva Kaul · 2020
Earlier work this paper cites.
“Linformer: Self-attention with Linear Complexity”
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang and Hao Ma · 2020
Earlier work this paper cites.
“Rethinking Attention with Performers”
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin and Lukasz Kaiser · 2021
Earlier work this paper cites.
“A Framework for Few-shot Language Model Evaluation”
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang and Andy Zou · 2021
Earlier work this paper cites.
“Combining Recurrent, Convolutional, and Continuous-time Models with the Linear State Space Layer”
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra and Christopher Ré · 2021
Earlier work this paper cites.
“Fnet: Mixing tokens with fourier transforms”
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein and Santiago Ontanon · 2021
Cited alongside, same era.
“When Attention Meets Fast Recurrence: Training Language Models with Reduced Compute”
Tao Lei · 2021
Cited alongside, same era.
“Random Feature Attention”
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith and Lingpeng Kong · 2021
Cited alongside, same era.
“Winogrande: An Adversarial Winograd Schema Challenge at Scale”
Keisuke Sakaguchi, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 2021
Cited alongside, same era.
“Linear Transformers are Secretly Fast Weight Programmers”
Imanol Schlag, Kazuki Irie and Jürgen Schmidhuber · 2021
Cited alongside, same era.
“NormFormer: Improved Transformer Pretraining with Extra Normalization”
“Resurrecting Recurrent Neural Networks for Long Sequences”
Antonio Orvieto, Samuel Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu and Soham De · 2023
Later among the works it cites.
“RWKV: Reinventing RNNs for the Transformer Era”
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella and Kranthi GV · 2023
Later among the works it cites.
“Exact computations with quasiseparable matrices”
Clément Pernet, Hippolyte Signargout and Gilles Villard · 2023
Later among the works it cites.
“Hyena Hierarchy: Towards Larger Convolutional Language Models”
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon and Christopher Ré · 2023
Later among the works it cites.
“Toeplitz Neural Network for Sequence Modeling”
Zhen Qin, Xiaodong Han, Weixuan Sun, Bowen He, Dong Li, Dongxu Li, Yuchao Dai, Lingpeng Kong and Yiran Zhong · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sam Shleifer, Jason Weston and Myle Ott · 2021
Cited alongside, same era.
“Roformer: Enhanced Transformer with Rotary Position Embedding”
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen and Yunfeng Liu · 2021
Cited alongside, same era.
“MLP-Mixer: An All-MLP Architecture for Vision”
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers and Jakob Uszkoreit · 2021
Cited alongside, same era.
“Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention”
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li and Vikas Singh · 2021
Cited alongside, same era.
“An Attention Free Transformer”
Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang and Josh Susskind · 2021
Cited alongside, same era.
“Gpt-NeoX-20B: An Open-source Autoregressive Language Model”
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell and Jason Phang · 2022
Cited alongside, same era.
“Monarch: Expressive structured matrices for efficient and accurate training”
Tri Dao, Beidi Chen, Nimit Sohoni, Arjun Desai, Michael Poli, Jessica Grogan, Alexander Liu, Aniruddh Rao, Atri Rudra and Christopher Ré · 2022
Cited alongside, same era.
Later among the works it cites.
“TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer”
Zhen Qin, Dong Li, Weigao Sun, Weixuan Sun, Xuyang Shen, Xiaodong Han, Yunshen Wei, Baohong Lv, Xiao Luo and Yu Qiao · 2023
Later among the works it cites.
“Hierarchically Gated Recurrent Neural Network for Sequence Modeling”
Zhen Qin, Songlin Yang and Yiran Zhong · 2023
Later among the works it cites.
“Simplified State Space Layers for Sequence Modeling”
Jimmy Smith, Andrew Warrington and Scott Linderman · 2023
Later among the works it cites.
“Retentive network: A successor to transformer for large language models”
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang and Furu Wei · 2023
Later among the works it cites.
“Llama: Open and Efficient Foundation Language Models”
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro and Faisal Azhar · 2023
Later among the works it cites.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava and Shruti Bhosale · 2023
Later among the works it cites.
Shida Wang and Beichen Xue · 2023
Later among the works it cites.
“Bytetransformer: A high-performance transformer boosted for variable-length inputs”
Yujia Zhai, Chengquan Jiang, Leyuan Wang, Xiaoying Jia, Shang Zhang, Zizhong Chen, Xin Liu and Yibo Zhu · 2023
Later among the works it cites.
“Linear Transformers with Learnable Kernel Functions are Better In-Context Models”
Yaroslav Aksenov, Nikita Balagansky, Sofia Vaina, Boris Shaposhnikov, Alexey Gorbatovski and Daniil Gavrilov · 2024
Closest in time.
“In-Context Language Learning: Architectures and Algorithms”
Ekin Akyürek, Bailin Wang, Yoon Kim and Jacob Andreas · 2024
Closest in time.
“The Hidden Attention of Mamba Models”, 2024
Ameen Ali, Itamar Zimerman and Lior Wolf · 2024
Closest in time.
“Zoology: Measuring and Improving Recall in Efficient Language Models”
Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra and Christopher Ré · 2024
Closest in time.
“Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff”
Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra and Christopher Ré · 2024
Closest in time.
“xLSTM: Extended Long Short-Term Memory”
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter and Sepp Hochreiter · 2024
Closest in time.
“RecurrentGemma: Moving Past Transformers for Efficient Open Language Models”
Aleksandar Botev, Soham De, Samuel Smith, Anushan Fernando, George-Cristian Muraru, Ruba Haroun, Leonard Berrada, Razvan Pascanu, Pier Sessa and Robert Dadashi · 2024
Closest in time.
“FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning”
Tri Dao · 2024
Closest in time.
“Vision Transformers Need Registers”
Timothée Darcet, Maxime Oquab, Julien Mairal and Piotr Bojanowski · 2024
Closest in time.
“Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models”
Soham De, Samuel Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen and Srivatsan Srinivasan · 2024
Closest in time.
“Fewer truncations improve language modeling”
Hantian Ding, Zijian Wang, Giovanni Paolini, Varun Kumar, Anoop Deoras, Dan Roth and Stefano Soatto · 2024
Closest in time.
“Monarch mixer: A simple sub-quadratic gemm-based architecture”
Dan Fu, Simran Arora, Jessica Grogan, Isys Johnson, Evan Eyuboglu, Armin Thomas, Benjamin Spector, Michael Poli, Atri Rudra and Christopher Ré · 2024
Closest in time.
“Zamba: A Compact 7B SSM Hybrid Model”
Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim and Beren Millidge · 2024
Closest in time.
“Is Mamba Capable of In-Context Learning?”
Riccardo Grazzi, Julien Siems, Simon Schrodi, Thomas Brox and Frank Hutter · 2024
Closest in time.
“Repeat After Me: Transformers Are Better Than State Space Models at Copying”
Samy Jelassi, David Brandfonbrener, Sham Kakade and Eran Malach · 2024
Closest in time.
“Jamba: A Hybrid Transformer-Mamba Language Model”
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov and Shai Shalev-Shwartz · 2024
Closest in time.
“World Model on Million-Length Video And Language With RingAttention”
Hao Liu, Wilson Yan, Matei Zaharia and Pieter Abbeel · 2024
Closest in time.
“Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks”
Jongho Park, Jaeseung Park, Zheyang Xiong, Nayoung Lee, Jaewoong Cho, Samet Oymak, Kangwook Lee and Dimitris Papailiopoulos · 2024
Closest in time.
“Eagle and Finch: RWKV with matrix-valued states and dynamic recurrence”
Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Teddy Ferdinan, Haowen Hou and Przemysław Kazienko · 2024
Closest in time.
“Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum”
Hadi Pouransari, Chun-Liang Li, Jen-Hao Chang, Pavan Vasu, Cem Koc, Vaishaal Shankar and Oncel Tuzel · 2024
Closest in time.
“HGRN2: Gated Linear RNNs with State Expansion”
Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun and Yiran Zhong · 2024
Closest in time.
“Chameleon: Mixed-Modal Early-Fusion Foundation Models”
Chameleon Team · 2024
Closest in time.
“Efficient Streaming Language Models with Attention Sinks”
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han and Mike Lewis · 2024
Closest in time.
“Gated Linear Attention Transformers with Hardware-Efficient Training”
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda and Yoon Kim · 2024
Closest in time.
“The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry”
Michael Zhang, Kush Bhatia, Hermann Kumbong and Christopher Ré · 2024
Closest in time.