Fetching the paper…
Reading the bibliography…
State-space models have gained popularity in sequence modelling due to their simple and efficient network structures.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1941
Earlier work this paper cites.
On the representation of continuous functions of many variables by superposition of continuous functions of one variable and addition , volume 28, pages 55–59
A. N. Kolmogorov · 1963
Earlier work this paper cites.
ON THE STRUCTURE OF REPRESENTATIONS OF CONTINUOUS FUNCTIONS OF SEVERAL VARIABLES AS FINITE SUMS OF CONTINUOUS FUNCTIONS OF ONE VARIABLE
David A Sprecher · 1965
Earlier work this paper cites.
Analytical Foundations of Volterra Series
Stephen Boyd, L. O. Chua, and C. A. Desoer · 1984
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R. Barron · 1994
Earlier work this paper cites.
A learning result for continuous-time recurrent neural networks
Eduardo D. Sontag · 1998
Earlier work this paper cites.
A state-space method for language modeling
V. Siivola and A. Honkela · 2003
Earlier work this paper cites.
On a Constructive Proof of Kolmogorov’s Superposition Theorem
Jürgen Braun and Michael Griebel · 2009
Earlier work this paper cites.
Approximationstheorie: Tschebyscheffsche Approximation mit Anwendungen
Lothar Collatz and Werner Krabs · 2013
Cited alongside, same era.
The Power of Depth for Feedforward Neural Networks
Ronen Eldan and Ohad Shamir · 2016
Cited alongside, same era.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2018
Cited alongside, same era.
Parallelizing Linear Recurrent Neural Nets Over Sequence Length
Eric Martin and Chris Cundy · 2018
Cited alongside, same era.
HiPPO: Recurrent Memory with Optimal Polynomial Projections
Approximation and Optimization Theory for Linear Continuous-Time Recurrent Neural Networks
Zhong Li, Jiequn Han, Weinan E, and Qianxiao Li · 2022
Later among the works it cites.
CKConv: Continuous Kernel Convolution For Sequential Data
David W. Romero, Anna Kuzina, Erik J. Bekkers, Jakub Mikolaj Tomczak, and Mark Hoogendoorn · 2022
Later among the works it cites.
Simplified State Space Layers for Sequence Modeling
Jimmy T. H. Smith, Andrew Warrington, and Scott Linderman · 2023
Closest in time.
Hungry Hungry Hippos: Towards Language Modeling with State Space Models
Daniel Y. Fu, Tri Dao, Khaled Kamal Saab, Armin W. Thomas, Atri Rudra, and Christopher Re · 2023
Closest in time.
How to Train your HIPPO: State Space Models with Generalized Orthogonal Basis Projections
Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Re · 2023
Closest in time.
Hyena Hierarchy: Towards Larger Convolutional Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
Cited alongside, same era.
Long Range Arena : A Benchmark for Efficient Transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Cited alongside, same era.
Learning Recurrent Neural Net Models of Nonlinear Systems
Joshua Hanson, Maxim Raginsky, and Eduardo Sontag · 2021
Cited alongside, same era.
Junxiong Wang, Jing Nathan Yan, Albert Gu, and Alexander M Rush · 2022
Cited alongside, same era.
On the Parameterization and Initialization of Diagonal State Space Models
Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré · 2022
Cited alongside, same era.
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Re · 2023
Closest in time.
Inverse approximation theory for nonlinear recurrent neural networks
Shida Wang, Zhong Li, and Qianxiao Li · 2023
Closest in time.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De · 2023
Closest in time.
A Brief Survey on the Approximation Theory for Sequence Modelling
Haotian Jiang, Qianxiao Li, Zhong Li, and Shida Wang · 2048
Closest in time.