Fetching the paper…
Reading the bibliography…
Selective state-space models (SSMs) are an emerging alternative to the Transformer, offering the unique advantage of parallel training and sequential inference.
Randomized Positional Encodings Boost Length Generalization of Transformers
Ruoss, A.; Delétang, G.; Genewein, T.; Grau-Moya, J.; Csordás, R.; Bennani, M.; Legg, S.; and Veness, J. 2023 · 1903
Earlier work this paper cites.
Memory-Augmented Recurrent Neural Networks Can Learn Generalized Dyck Languages
Suzgun, M.; Gehrmann, S.; Belinkov, Y.; and Shieber, S. M. 2019 · 1911
Earlier work this paper cites.
Neural Nets and the Brain Model Problem
Minsky, M. 1954 · 1954
Earlier work this paper cites.
Algebraic Theory of Machines. I. Prime Decomposition Theorem for Finite Semigroups and Machines
Krohn, K.; and Rhodes, J. 1965 · 1965
Earlier work this paper cites.
Computation: finite and infinite machines
Minsky, M. L. 1967 · 1967
Earlier work this paper cites.
Threshold matrices and the state assignment problem for neural nets
Dewdney, A. K. 1977 · 1977
Earlier work this paper cites.
Prefix Sums and Their Applications
Blelloch, G. E. 1990 · 1990
Earlier work this paper cites.
Finding Structure in Time
Elman, J. L. 1990 · 1990
Earlier work this paper cites.
Finite automata, formal logic, and circuit complexity
Straubing, H. 1994 · 1994
Earlier work this paper cites.
Optimal simulation of automata by neural nets
Indyk, P. 1995 · 1995
Earlier work this paper cites.
On the Computational Power of Neural Nets
Siegelmann, H.; and Sontag, E. 1995 · 1995
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Visual group theory
Carter, N. C. 2009 · 2009
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, A.; and Mikolov, T. 2015 · 2015
Earlier work this paper cites.
Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Attention Is All You Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Parallelizing Linear Recurrent Neural Nets Over Sequence Length
Martin, E.; and Cundy, C. 2018 · 2018
Cited alongside, same era.
On the Ability and Limitations of Transformers to Recognize Formal Languages
Bhattamishra, S.; Ahuja, K.; and Goyal, N. 2020 · 2020
Cited alongside, same era.
HiPPO: Recurrent Memory with Optimal Polynomial Projections
Gu, A.; Dao, T.; Ermon, S.; Rudra, A.; and Re, C. 2020 · 2020
Cited alongside, same era.
Theoretical Limitations of Self-Attention in Neural Sequence Models
Hahn, M. 2020 · 2020
Cited alongside, same era.
Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers
Gu, A.; Johnson, I.; Goel, K.; Saab, K.; Dao, T.; Rudra, A.; and Ré, C. 2021 · 2021
Hyena Hierarchy: Towards Larger Convolutional Language Models
Poli, M.; Massaroli, S.; Nguyen, E.; Fu, D. Y.; Dao, T.; Baccus, S.; Bengio, Y.; Ermon, S.; and Ré, C. 2023 · 2023
Later among the works it cites.
Simplified State Space Layers for Sequence Modeling
Smith, J. T. H.; Warrington, A.; and Linderman, S. W. 2023 · 2023
Later among the works it cites.
Efficiently Representing Finite-state Automata With Recurrent Neural Networks
Svete, A.; and Cotterell, R. 2023 · 2023
Later among the works it cites.
Theoretical Foundations of Deep Selective State-Space Models
Cirone, N. M.; Orvieto, A.; Walker, B.; Salvi, C.; and Lyons, T. 2024 · 2024
Closest in time.
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Dao, T. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the Parameterization and Initialization of Diagonal State Space Models
Gu, A.; Goel, K.; Gupta, A.; and Ré, C. 2022 · 2022
Cited alongside, same era.
Efficiently Modeling Long Sequences with Structured State Spaces
Gu, A.; Goel, K.; and Ré, C. 2022 · 2022
Cited alongside, same era.
Diagonal State Spaces are as Effective as Structured State Spaces
Gupta, A.; Gu, A.; and Berant, J. 2022 · 2022
Cited alongside, same era.
Neural Networks and the Chomsky Hierarchy
Delétang, G.; Ruoss, A.; Grau-Moya, J.; Genewein, T.; Wenliang, L. K.; Catt, E.; Cundy, C.; Hutter, M.; Legg, S.; Veness, J.; and Ortega, P. A. 2023 · 2023
Cited alongside, same era.
Hungry Hungry Hippos: Towards Language Modeling with State Space Models
Fu, D. Y.; Dao, T.; Saab, K. K.; Thomas, A. W.; Rudra, A.; and Ré, C. 2023 · 2023
Cited alongside, same era.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Gu, A.; and Dao, T. 2023 · 2023
Cited alongside, same era.
Closest in time.
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
De, S.; Smith, S. L.; Fernando, A.; Botev, A.; Cristian-Muraru, G.; Gu, A.; Haroun, R.; Berrada, L.; Chen, Y.; Srinivasan, S.; Desjardins, G.; Doucet, A.; Budden, D.; Teh, Y. W.; Pascanu, R.; De Freitas, N.; and Gulcehre, C. 2024 · 2024
Closest in time.
Advancing Regular Language Reasoning in Linear Recurrent Neural Networks
Fan, T.-H.; Chi, T.-C.; and Rudnicky, A. I. 2024 · 2024
Closest in time.
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
Grazzi, R.; Siems, J.; Franke, J. K.; Zela, A.; Hutter, F.; and Pontil, M. 2024 · 2024
Closest in time.
The Illusion of State in State-Space Models
Merrill, W.; Petty, J.; and Sabharwal, A. 2024 · 2024
Closest in time.
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
Orvieto, A.; De, S.; Gulcehre, C.; Pascanu, R.; and Smith, S. L. 2024 · 2024
Closest in time.
The Expressive Capacity of State Space Models: A Formal Language Perspective
Sarrof, Y.; Veitsman, Y.; and Hahn, M. 2024 · 2024
Closest in time.
What Formal Languages Can Transformers Express? A Survey
Strobl, L.; Merrill, W.; Weiss, G.; Chiang, D.; and Angluin, D. 2024 · 2024
Closest in time.
On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages
Terzić, A.; Hersche, M.; Camposampiero, G.; Hofmann, T.; Sebastian, A.; and Rahimi, A. 2024 · 2024
Closest in time.
Limits of Deep Learning: Sequence Modeling through the Lens of Complexity Theory
Zubić, N.; Soldá, F.; Sulser, A.; and Scaramuzza, D. 2024 · 2024
Closest in time.