Fetching the paper…
Reading the bibliography…
State space models (SSMs) leverage linear, time-invariant (LTI) systems to effectively learn sequences with long-range dependencies.
Analytic properties of schmidt pairs for a hankel operator and the generalized schur–takagi problem
Vadim Movsesovich Adamyan, Damir Zyamovich Arov, and Mark Grigor’evich Krein · 1971
Earlier work this paper cites.
All optimal hankel-norm approximations of linear multivariable systems and their L, ∞ \infty -error bounds
Keith Glover · 1984
Earlier work this paper cites.
Density-Functional Theory of Atoms and Molecules
Robert G. Parr and Weitao Yang · 1994
Earlier work this paper cites.
Discrete-time signal processing
Alan V Oppenheim · 1999
Earlier work this paper cites.
Calculus of Variations
Izrail Moiseevitch Gelfand, Richard A Silverman, et al · 2000
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
The MNIST database of handwritten digit images for machine learning research
Li Deng · 2012
Earlier work this paper cites.
Approximate computation and implicit regularization for very large-scale data analysis
M. W. Mahoney · 2012
Earlier work this paper cites.
Field quantization
Walter Greiner and Joachim Reinhardt · 2013
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Earlier work this paper cites.
Essential radio astronomy
James J Condon and Scott M Ransom · 2016
Earlier work this paper cites.
Sobolev training for neural networks
W. M. Czarnecki, S. Osindero, M. Jaderberg, G. Swirszcz, and R. Pascanu · 2017
Earlier work this paper cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2018
Earlier work this paper cites.
Gradient descent learns linear dynamical systems
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
S. Arora, S. S. Du, W. Hu, Z. Li, and R. Wang · 2019
Earlier work this paper cites.
The convergence rate of neural networks for learned functions of different frequencies
R. Basri, D. Jacobs, Y. Kasten, and S. Kritchman · 2019
Earlier work this paper cites.
On the inductive bias of neural tangent kernels
A. Bietti and J. Mairal · 2019
Earlier work this paper cites.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2019
Earlier work this paper cites.
Antisymmetricrnn: A dynamical system view on recurrent neural networks
Bo Chang, Minmin Chen, Eldad Haber, and Ed H Chi · 2019
Earlier work this paper cites.
On the spectral bias of neural networks
N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. Hamprecht, Y. Bengio, and A. Courville · 2019
Earlier work this paper cites.
On learning over-parameterized neural networks: A functional approximation perspective
L. Su and P. Yang · 2019
Earlier work this paper cites.
Approximation theory and approximation practice, extended edition
Lloyd N Trefethen · 2019
Cited alongside, same era.
Legendre memory units: Continuous-time representation in recurrent neural networks
Aaron Voelker, Ivana Kajić, and Chris Eliasmith · 2019
Cited alongside, same era.
A fine-grained spectral perspective on neural networks
G. Yang and H. Salman · 2019
Cited alongside, same era.
Frequency bias in neural networks for input of non-uniform density
R. Basri, M. Galun, A. Geifman, D. Jacobs, Y. Kasten, and S. Kritchman · 2020
Cited alongside, same era.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Cited alongside, same era.
Hippo: Recurrent memory with optimal polynomial projections
Ckconv: Continuous kernel convolution for sequential data
David W Romero, Anna Kuzina, Erik J Bekkers, Jakub M Tomczak, and Mark Hoogendoorn · 2022
Later among the works it cites.
Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting
Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin · 2022
Later among the works it cites.
Naman Agarwal, Daniel Suo, Xinyi Chen, and Elad Hazan · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Later among the works it cites.
How to train your hippo: State space models with generalized orthogonal basis projections
Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Ré · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Introduction to probability
Dennis Sun · 2020
Cited alongside, same era.
Frequency principle: Fourier analysis sheds light on deep neural networks
Z.-Q. J. Xu · 2020
Cited alongside, same era.
Lipschitz recurrent neural networks
N Benjamin Erichson, Omri Azencot, Alejandro Queiruga, Liam Hodgkinson, and Michael W Mahoney · 2021
Cited alongside, same era.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Cited alongside, same era.
Liquid structural state-space models
Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, and Daniela Rus · 2023
Later among the works it cites.
A time series is worth 64 words: Long-term forecasting with transformers
Yuqi Nie, Nam H Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam · 2023
Later among the works it cites.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
Jimmy T.H. Smith, Andrew Warrington, and Scott Linderman · 2023
Later among the works it cites.
Sobolev acceleration for neural networks
Hwijae Son · 2023
Later among the works it cites.
Stablessm: Alleviating the curse of memory in state-space models through stable reparameterization
Shida Wang and Qianxiao Li · 2023
Later among the works it cites.
Tuning frequency bias in neural network training with nonuniform data
Annan Yu, Yunan Yang, and Alex Townsend · 2023
Later among the works it cites.
Tri Dao and Albert Gu · 2024
Closest in time.
From generalization analysis to optimization designs for state space models
Fusheng Liu and Qianxiao Li · 2024
Closest in time.
The role of state matrix initialization in ssms: A perspective on the approximation-estimation tradeoff
Fusheng Liu and Qianxiao Li · 2024
Closest in time.
Mitigating spectral bias for the multiscale operator learning
Xinliang Liu, Bo Xu, Shuhao Cao, and Lei Zhang · 2024
Closest in time.
Universality of linear recurrences followed by non-linear projections: Finite-width guarantees and benefits of complex eigenvalues
Antonio Orvieto, Soham De, Caglar Gulcehre, Razvan Pascanu, and Samuel L Smith · 2024
Closest in time.
S4++: Elevating long sequence modeling with state memory reply
Biqing Qi, Junqi Gao, Dong Li, Kaiyan Zhang, Jianxing Liu, Ligang Wu, and Bowen Zhou · 2024
Closest in time.
Towards a theory of learning dynamics in deep state space models
Jakub Smékal, Jimmy TH Smith, Michael Kleinman, Dan Biderman, and Scott W Linderman · 2024
Closest in time.
Convolutional state space models for long-range spatiotemporal modeling
Jimmy Smith, Shalini De Mello, Jan Kautz, Scott Linderman, and Wonmin Byeon · 2024
Closest in time.
State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory
Shida Wang and Beichen Xue · 2024
Closest in time.
There is hope to avoid hippos for long-memory state space models
Annan Yu, Michael W Mahoney, and N Benjamin Erichson · 2024
Closest in time.
Robustifying state-space models for long sequences via approximate diagonalization
Annan Yu, Arnur Nigmetov, Dmitriy Morozov, Michael W. Mahoney, and N. Benjamin Erichson · 2024
Closest in time.
Leveraging the hankel norm approximation and data-driven algorithms in reduced order modeling
Annan Yu and Alex Townsend · 2024
Closest in time.