Fetching the paper…
Reading the bibliography…
In recent years, there has been a growing interest in integrating linear state-space models (SSM) in deep neural network architectures of foundation models.
1912
Earlier work this paper cites.
J. Sherman and W. J. Morrison, “Adjustment of an Inverse Matrix Corresponding to a Change in One Element of a Given Matrix,” The Annals of Mathematical Statistics , vol. 21, no. 1, pp. 124 – 127, 1950
1950
Earlier work this paper cites.
G. E. Blelloch, “Prefix sums and their applications,” 1990
1990
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
2004
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the thirteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 2010, pp. 249–256
2010
Earlier work this paper cites.
2014
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning . MIT press, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems , vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter, “Self-normalizing neural networks,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang et al. , “Big bird: Transformers for longer sequences,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Cited alongside, same era.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The Efficient Transformer,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré, “HiPPO: Recurrent Memory with Optimal Polynomial Projections,” in Advances in Neural Information Processing Systems , vol. 33. Curran Associates, Inc., 2020, pp. 1474–1487
2020
Cited alongside, same era.
Y. Tay, M. Dehghani, S. Abnar, Y. Shen, D. Bahri, P. Pham, J. Rao, L. Yang, S. Ruder, and D. Metzler, “Long Range Arena : A Benchmark for Efficient Transformers,” in International Conference on Learning Representations (ICLR) , 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Gupta, A. Gu, and J. Berant, “Diagonal state spaces are as effective as structured state spaces,” in Advances in Neural Information Processing Systems , vol. 35. Curran Associates, Inc., 2022, pp. 22 982–22 994
2022
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?” in Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , 2021, pp. 610–623
2021
Cited alongside, same era.
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller, “Rethinking Attention with Performers,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
A. Gu, I. Johnson, K. Goel, K. Saab, T. Dao, A. Rudra, and C. Ré, “Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space Layers,” in Advances in Neural Information Processing Systems , vol. 34. Curran Associates, Inc., 2021, pp. 572–585
2021
Cited alongside, same era.
H. Liu, Z. Dai, D. So, and Q. V. Le, “Pay Attention to MLPs,” in Advances in Neural Information Processing Systems , A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., 2021
2021
Cited alongside, same era.
A. Gu, K. Goel, and C. Ré, “Efficiently Modeling Long Sequences with Structured State Spaces,” in The International Conference on Learning Representations (ICLR) , 2022
2022
Cited alongside, same era.
J. T. Smith, A. Warrington, and S. Linderman, “Simplified State Space Layers for Sequence Modeling,” in The Eleventh International Conference on Learning Representations , 2023
2023
Later among the works it cites.
A. Orvieto, S. L. Smith, A. Gu, A. Fernando, C. Gulcehre, R. Pascanu, and S. De, “Resurrecting Recurrent Neural Networks for Long Sequences,” in Proceedings of the 40th International Conference on Machine Learning , vol. 202. PMLR, 23–29 Jul 2023, pp. 26 670–26 698
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.