Fetching the paper…
Reading the bibliography…
In this paper, we investigate the length-extension of state-space models (SSMs) in language modeling.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1941
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Long Short-term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The Vanishing Gradient Problem During Learning Recurrent Neural Nets and Problem Solutions
Sepp Hochreiter · 1998
Earlier work this paper cites.
Tutorial on training recurrent neural networks, covering BPPT, RTRL, EKF and the echo state network approach
Herbert Jaeger · 2002
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Parallelizing Linear Recurrent Neural Nets Over Sequence Length
Eric Martin and Chris Cundy · 2018
Earlier work this paper cites.
Neural Ordinary Differential Equations, December 2019
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud · 2019
Earlier work this paper cites.
Stable Extrapolation of Analytic Functions
Laurent Demanet and Alex Townsend · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, and Amanda Askell · 2020
Cited alongside, same era.
HiPPO: Recurrent Memory with Optimal Polynomial Projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
Cited alongside, same era.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, August 2020
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
A unified framework of online learning algorithms for training recurrent neural networks
Owen Marschall, Kyunghyun Cho, and Cristina Savin · 2020
Cited alongside, same era.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling, December 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Wordcraft: Story Writing With Large Language Models
Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito · 2022
Later among the works it cites.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces, December 2023
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei · 2023
Later among the works it cites.
Gated Linear Attention Transformers with Hardware-Efficient Training, December 2023
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim · 2023
Later among the works it cites.
State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory
Shida Wang and Beichen Xue · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
Zhong Li, Jiequn Han, Weinan E, and Qianxiao Li · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Efficiently Modeling Long Sequences with Structured State Spaces
Albert Gu, Karan Goel, and Christopher Re · 2021
Cited alongside, same era.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation, April 2022
Ofir Press, Noah A. Smith, and Mike Lewis · 2022
Cited alongside, same era.
RoFormer: Enhanced Transformer with Rotary Position Embedding, August 2022
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2022
Cited alongside, same era.
A Length-Extrapolatable Transformer, December 2022
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei · 2022
Cited alongside, same era.
CLEX: Continuous Length Extrapolation for Large Language Models, October 2023a
Guanzheng Chen, Xin Li, Zaiqiao Meng, Shangsong Liang, and Lidong Bing
Cited in the paper.
Simplified State Space Layers for Sequence Modeling
Jimmy T. H. Smith, Andrew Warrington, and Scott Linderman · 2023
Later among the works it cites.
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
Shida Wang, Zhong Li, and Qianxiao Li · 2023
Later among the works it cites.
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization, November 2023
Shida Wang and Qianxiao Li · 2023
Later among the works it cites.
Position Interpolation Improves ALiBi Extrapolation, October 2023
Faisal Al-Khateeb, Nolan Dey, Daria Soboleva, and Joel Hestness · 2023
Later among the works it cites.
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models, February 2024
Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, Arnaud Doucet, David Budden, Yee Whye Teh, Razvan Pascanu, Nando De Freitas, and Caglar Gulcehre · 2024
Closest in time.
A Brief Survey on the Approximation Theory for Sequence Modelling
Haotian Jiang, Qianxiao Li, Zhong Li, and Shida Wang · 2048
Closest in time.