Fetching the paper…
Reading the bibliography…
In this paper, we investigate the long-term memory learning capabilities of state-space models (SSMs) from the perspective of parameterization.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., and Frasconi, P · 1941
Earlier work this paper cites.
Analytical Foundations of Volterra Series
Boyd, S., Chua, L. O., and Desoer, C. A · 1984
Earlier work this paper cites.
Fading memory and the problem of approximating nonlinear operators with Volterra series
Boyd, S. and Chua, L · 1985
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Algorithms for Discrete Fourier Transform and Convolution
Tolimieri, R., An, M., and Lu, C · 1989
Earlier work this paper cites.
Recurrent neural networks and robust time series prediction
Connor, J. T., Martin, R. D., and Atlas, L. E · 1994
Earlier work this paper cites.
Long Short-term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
The Vanishing Gradient Problem During Learning Recurrent Neural Nets and Problem Solutions
Hochreiter, S · 1998
Earlier work this paper cites.
A state-space method for language modeling
Siivola, V. and Honkela, A · 2003
Earlier work this paper cites.
Princeton Lectures in Analysis
Stein, E. M. and Shakarchi, R · 2003
Earlier work this paper cites.
Generating Text with Recurrent Neural Networks
Sutskever, I., Martens, J., and Hinton, G · 2011
Earlier work this paper cites.
Pointer Sentinel Mixture Models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Cited alongside, same era.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Parallelizing Linear Recurrent Neural Nets Over Sequence Length
Martin, E. and Cundy, C · 2018
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Li, Y., Wei, C., and Ma, T · 2019
Cited alongside, same era.
Super-convergence: Very fast training of neural networks using large learning rates
Smith, L. N. and Topin, N · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A · 2020
Coupled Oscillatory Recurrent Neural Network (coRNN): An accurate and (gradient) stable architecture for learning long time dependencies
Rusch, T. K. and Mishra, S · 2022
Later among the works it cites.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces, December 2023
Gu, A. and Dao, T · 2023
Closest in time.
RWKV: Reinventing RNNs for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al · 2023
Closest in time.
Hyena Hierarchy: Towards Larger Convolutional Language Models
Poli, M., Massaroli, S., Nguyen, E., Fu, D. Y., Dao, T., Baccus, S., Bengio, Y., Ermon, S., and Re, C · 2023
Closest in time.
Simplified State Space Layers for Sequence Modeling
Smith, J. T. H., Warrington, A., and Linderman, S · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
HiPPO: Recurrent Memory with Optimal Polynomial Projections
Gu, A., Dao, T., Ermon, S., Rudra, A., and Ré, C · 2020
Cited alongside, same era.
On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
Li, Z., Han, J., E, W., and Li, Q · 2020
Cited alongside, same era.
Long Range Arena : A Benchmark for Efficient Transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Cited alongside, same era.
Approximation and Optimization Theory for Linear Continuous-Time Recurrent Neural Networks
Li, Z., Han, J., E, W., and Li, Q · 2022
Cited alongside, same era.
On the Parameterization and Initialization of Diagonal State Space Models
Gu, A., Goel, K., Gupta, A., and Ré, C
Cited in the paper.
Efficiently Modeling Long Sequences with Structured State Spaces
Gu, A., Goel, K., and Re, C
Cited in the paper.
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F · 2023
Closest in time.
State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory
Wang, S. and Xue, B · 2023
Closest in time.
Improve long-term memory learning through rescaling the error temporally
Wang, S. and Yan, Z · 2023
Closest in time.
Inverse Approximation Theory for Nonlinear Recurrent Neural Networks
Wang, S., Li, Z., and Li, Q · 2023
Closest in time.
A Brief Survey on the Approximation Theory for Sequence Modelling
Jiang, H., Li, Q., Li, Z., and Wang, S · 2048
Closest in time.