Fetching the paper…
Reading the bibliography…
Linear recurrent neural networks (RNNs) and state-space models (SSMs) such as Mamba have become promising alternatives to softmax-attention as sequence mixing layers in Transformer architectures.
Ghaoui, L. E., Gu, F., Travacca, B., Askari, A., and Tsai, A. Y · 1908
Earlier work this paper cites.
Bai, S., Kolter, J. Z., and Koltun, V · 1909
Earlier work this paper cites.
Sur les opérations dans les ensembles abstraits et leur application aux équations intégrales
Banach, S · 1922
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
Hopfield, J. J · 1982
Earlier work this paper cites.
Sequential thought processes in pdp models
Rumelhart, D. E., Smolensky, P., McClelland, J. L., and Hinton, G · 1986
Earlier work this paper cites.
Finding structure in time
Elman, J. L · 1990
Earlier work this paper cites.
On the computational power of neural nets
Siegelmann, H. T. and Sontag, E. D · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Hochreiter, S., Bengio, Y., Frasconi, P., et al · 2001
Earlier work this paper cites.
The "echo state" approach to analysing and training recurrent neural networks-with an erratum note
Jaeger, H · 2001
Earlier work this paper cites.
Fixed point theory , volume 14
Granas, A., Dugundji, J., et al · 2003
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity, 2020
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2006
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., et al · 2009
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2011
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
Towards AI-complete question answering: A set of prerequisite toy tasks, 2015
Weston, J., Bordes, A., Chopra, S., Rush, A. M., van Merriënboer, B., Joulin, A., and Mikolov, T · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Arjovsky, M., Shah, A., and Bengio, Y · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks, 2016
Graves, A · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Neural ordinary differential equations
Chen, R. T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K · 2018
Earlier work this paper cites.
Orthogonal recurrent neural networks with scaled Cayley transform
Helfrich, K., Willmott, D., and Ye, Q · 2018
Cited alongside, same era.
Reviving and improving recurrent back-propagation
Liao, R., Xiong, Y., Fetaya, E., Zhang, L., Yoon, K., Pitkow, X., Urtasun, R., and Zemel, R · 2018
Cited alongside, same era.
Parallelizing linear recurrent neural nets over sequence length
Martin, E. and Cundy, C · 2018
Cited alongside, same era.
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L · 2019
Cited alongside, same era.
On the computational power of RNNs
Korsky, S. A · 2019
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2024
Later among the works it cites.
Repeat after me: Transformers are better than state space models at copying
Jelassi, S., Brandfonbrener, D., Kakade, S. M., and Malach, E · 2024
Later among the works it cites.
Parallelizing non-linear sequential models over the sequence length
Lim, Y. H., Zhu, Q., Selfridge, J., and Kasim, M. F · 2024
Later among the works it cites.
Vmamba: Visual state space model
Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., Jiao, J., and Liu, Y · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Universal simulation of stable dynamical systems by recurrent neural nets
Hanson, J. and Raginsky, M · 2020
Cited alongside, same era.
Stabilizing equilibrium models by jacobian regularization
Bai, S., Koltun, V., and Kolter, J. Z · 2021
Cited alongside, same era.
Skyformer: Remodel self-attention with Gaussian kernel and Nystrom method
Chen, Y., Zeng, Q., Ji, H., and Yang, Y · 2021
Cited alongside, same era.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
Schwarzschild, A., Borgnia, E., Gupta, A., Huang, F., Vishkin, U., Goldblum, M., and Goldstein, T · 2021
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Ré, C · 2022
Cited alongside, same era.
Looped transformers as programmable computers
Giannou, A., Rajput, S., Sohn, J., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Cited alongside, same era.
The parallelism tradeoff: Limitations of log-precision transformers
Merrill, W. and Sabharwal, A · 2023
Cited alongside, same era.
Merrill, W., Petty, J., and Sabharwal, A · 2024
Later among the works it cites.
Sequence modeling and design from molecular to genome scale with Evo
Nguyen, E., Poli, M., Durrant, M. G., Kang, B., Katrekar, D., Li, D. B., Bartie, L. J., Thomas, A. W., King, S. H., Brixi, G., et al · 2024
Later among the works it cites.
The fineweb datasets: Decanting the web for the finest text data at scale
Penedo, G., Kydlíček, H., allal, L. B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L. V., and Wolf, T · 2024
Later among the works it cites.
Eagle and Finch: RWKV with matrix-valued states and dynamic recurrence
Peng, B., Goldstein, D., Anthony, Q., Albalak, A., Alcaide, E., Biderman, S., Cheah, E., Du, X., Ferdinan, T., Hou, H., et al · 2024
Later among the works it cites.
HGRN2: Gated linear RNNs with state expansion
Qin, Z., Yang, S., Sun, W., Shen, X., Li, D., Sun, W., and Zhong, Y · 2024
Later among the works it cites.
The expressive capacity of state space models: A formal language perspective
Sarrof, Y., Veitsman, Y., and Hahn, M · 2024
Later among the works it cites.
Mimetic initialization helps state space models learn to recall
Trockman, A., Harutyunyan, H., Kolter, J. Z., Kumar, S., and Bhojanapalli, S · 2024
Later among the works it cites.
An empirical study of Mamba-based language models, 2024
Waleffe, R., Byeon, W., Riach, D., Norick, B., Korthikanti, V., Dao, T., Gu, A., Hatamizadeh, A., Singh, S., Narayanan, D., et al · 2024
Later among the works it cites.
Gated linear attention transformers with hardware-efficient training
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2024
Later among the works it cites.
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
Zucchet, N. and Orvieto, A · 2024
Later among the works it cites.
Artificial kuramoto oscillatory neurons
Miyato, T., Löwe, S., Geiger, A., and Welling, M · 2025
Closest in time.
Rwkv-7 "goose" with expressive dynamic state evolution
Peng, B., Zhang, R., Goldstein, D., Alcaide, E., Du, X., Hou, H., Lin, J., Liu, J., Lu, J., Merrill, W., Song, G., Tan, K., Utpala, S., Wilce, N., Wind, J. S., Wu, T., Wuttke, D., and Zhou-Zheng, C · 2025
Closest in time.
Reasoning with latent thoughts: On the power of looped transformers
Saunshi, N., Dikkala, N., Li, Z., Kumar, S., and Reddi, S. J · 2025
Closest in time.
Implicit language models are rnns: Balancing parallelization and expressivity
Schöne, M., Rahmani, B., Kremer, H., Falck, F., Ballani, H., and Gladrow, J · 2025
Closest in time.
Deltaproduct: Increasing the expressivity of deltanet through products of householders
Siems, J., Carstensen, T., Zela, A., Hutter, F., Pontil, M., and Grazzi, R · 2025
Closest in time.
On the expressiveness and length generalization of selective state space models on regular languages
Terzic, A., Hersche, M., Camposampiero, G., Hofmann, T., Sebastian, A., and Rahimi, A · 2025
Closest in time.
Gated delta networks: Improving mamba2 with delta rule
Yang, S., Kautz, J., and Hatamizadeh, A · 2025
Closest in time.