Fetching the paper…
Reading the bibliography…
Modern neural sequence models are designed to meet the dual mandate of parallelizable training and fast sequential inference.
Parallelism in comparison problems
Leslie G Valiant · 1975
Earlier work this paper cites.
Bounded-width polynomial-size branching programs recognize exactly those languages in nc1
David A. Mix Barrington · 1986
Earlier work this paper cites.
Attractor dynamics and parallelism in a connectionist sequential machine
Michael I Jordan · 1986
Earlier work this paper cites.
A focused backpropagation algorithm for temporal pattern recognition
Michael C. Mozer · 1989
Earlier work this paper cites.
Prefix sums and their applications
Guy E Blelloch · 1990
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Çağlar Gülçehre, Bart van Merriënboer, Dzmitry Bahdanau, Fethi Bougares Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Strongly-typed recurrent neural networks
David Balduzzi and Muhammad Ghifary · 2016
Earlier work this paper cites.
Neural machine translation in linear time
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Quasi-recurrent neural networks
James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Simple recurrent units for highly parallelizable recurrence
Tao Lei, Yu Zhang, Sida I. Wang, Hui Dai, and Yoav Artzi · 2018
Cited alongside, same era.
Independently recurrent neural network (IndRNN): Building a longer and deeper RNN
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao · 2018
Cited alongside, same era.
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy · 2018
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke et al · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Hierarchically gated recurrent neural network for sequence modeling
Zhen Qin, Songlin Yang, and Yiran Zhong · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei · 2023
Later among the works it cites.
Simple linear attention language models balance the recall-throughput tradeoff
Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré · 2024
Later among the works it cites.
xLSTM: Extended long short-term memory
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2024
Later among the works it cites.
Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality
Tri Dao and Albert Gu · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Hahn · 2020
Cited alongside, same era.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
A formal hierarchy of RNN architectures
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz, Noah A. Smith, and Eran Yahav · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Cited alongside, same era.
Going beyond linear transformers with recurrent fast weight programmers
Kazuki Irie, Imanol Schlag, Róbert Csordás, and Jürgen Schmidhuber · 2021
Cited alongside, same era.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A Smith, and Lingpeng Kong · 2021
Cited alongside, same era.
Linear Transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber · 2021
Cited alongside, same era.
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2024
Later among the works it cites.
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal · 2024
Later among the works it cites.
What formal languages can transformers express? a survey
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin · 2024
Later among the works it cites.
Unlocking state-tracking in linear rnns through negative eigenvalues
Riccardo Grazzi, Julien Siems, Jörg KH Franke, Arber Zela, Frank Hutter, and Massimiliano Pontil · 2025
Closest in time.
Han Guo, Songlin Yang, Tarushii Goel, Eric P Xing, Tri Dao, and Yoon Kim · 2025
Closest in time.
(How) do language models track state?
Belinda Z Li, Zifan Carl Guo, and Jacob Andreas · 2025
Closest in time.
Fixed-point RNNs: From diagonal to dense in a few iterations
Sajad Movahedi, Felix Sarnthein, Nicola Muca Cirone, and Antonio Orvieto · 2025
Closest in time.
Revisiting associative recall in modern recurrent models
Destiny Okpekpe and Antonio Orvieto · 2025
Closest in time.
RWKV-7" goose" with expressive dynamic state evolution
Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Xingjian Du, Haowen Hou, Jiaju Lin, Jiaxing Liu, Janna Lu, William Merrill, et al · 2025
Closest in time.
Deltaproduct: Improving state-tracking in linear rnns via householder products
Julien Siems, Timur Carstensen, Arber Zela, Frank Hutter, Massimiliano Pontil, and Riccardo Grazzi · 2025
Closest in time.
Gated delta networks: Improving Mamba2 with delta rule
Songlin Yang, Jan Kautz, and Ali Hatamizadeh · 2025
Closest in time.