Fetching the paper…
Reading the bibliography…
Linear Recurrent Neural Networks (linear RNNs) have emerged as competitive alternatives to Transformers for sequence modeling, offering efficient training and linear-time inference.
LIII. On lines and planes of closest fit to systems of points in space
Karl Pearson · 1901
Earlier work this paper cites.
Unitary triangularization of a nonsymmetric matrix
Alston S Householder · 1958
Earlier work this paper cites.
Linear automaton transformations
Anil Nerode · 1958
Earlier work this paper cites.
Algebraic theory of machines. i. prime decomposition theorem for finite semigroups and machines
Kenneth Krohn and John Rhodes · 1965
Earlier work this paper cites.
Block reflectors: Theory and computation
Robert Schreiber and Beresford Parlett · 1988
Earlier work this paper cites.
Prefix sums and their applications
Guy E. Blelloch · 1990
Earlier work this paper cites.
On the symmetry group of the dodecahedron
Lorraine L Foster · 1990
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen netzen
Sepp Hochreiter · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Olivier Roy and Martin Vetterli · 2007
Earlier work this paper cites.
Sampling from large matrices: An approach through geometric functional analysis
Mark Rudelson and Roman Vershynin · 2007
Earlier work this paper cites.
Groups and symmetries from finite groups to lie groups, 2010
Yvette Kosmann Schwarzbach · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
DizzyRNN: Reparameterizing Recurrent Neural Networks for Norm-Preserving Backpropagation
Victor D. Dorobantu, Per Andre Stromhaug, and Jess Renteria · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Tunable Efficient Unitary Neural Networks (EUNN) and their application to RNNs
L. Jing, Y. Shen, T. Dubcek, J. Peurifoy, S. Skirlo, Y. LeCun, M. Tegmark, and M. Soljacic · 2017
Earlier work this paper cites.
Efficient orthogonal parametrisation of recurrent neural networks using householder reflections
Z. Mhammedi, A. Hellicar, A. Rahman, and J. Bailey · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Parallelizing Linear Recurrent Neural Nets Over Sequence Length
E. Martin and C. Cundy · 2018
Earlier work this paper cites.
Kronecker recurrent units
C. Jose, M. Cisse, and F. Fleuret · 2018
Earlier work this paper cites.
Think you have solved question answering? Try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Universal Transformers
D. Mostafa, G. Stephan, V. Oriol, J. Uszkoreit, and L. Kaiser · 2019
Earlier work this paper cites.
Rotational unit of memory: a novel representation unit for rnns with scalable applications
Rumen Dangovski, Li Jing, Preslav Nakov, Mićo Tatalović, and Marin Soljačić · 2019
Cited alongside, same era.
Triton: An intermediate language and compiler for tiled neural network computations
Philippe Tillet, Hsiang-Tsung Kung, and David Cox · 2019
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Cited alongside, same era.
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
xLSTM: Extended Long Short-Term Memory
M. Beck, K. Pöppel, M. Spanring, A. Auer, O. Prudnikova, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter · 2024
Later among the works it cites.
Learning to (learn at test time): RNNs with expressive hidden states
Y. Sun, X. Li, K. Dalal, J. Xu, A. Vikram, G. Zhang, Y. Dubois, X. Chen, X. Wang, S. Koyejo, T. Hashimoto, and C. Guestrin · 2024
Later among the works it cites.
Titans: Learning to memorize at test time
Ali Behrouz, Peilin Zhong, and Vahab Mirrokni · 2024
Later among the works it cites.
The Expressive Capacity of State Space Models: A Formal Language Perspective
Y. Sarrof, Y. Veitsman, and M. Hahn · 2024
Later among the works it cites.
Theoretical Foundations of Deep Selective State-Space Models
N. M. Cirone, A. Orvieto, B. Walker, C. Salvi, and T. Lyons · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Hahn · 2020
Cited alongside, same era.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
PIQA: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le bras, Jianfeng Gao, and Yejin Choi · 2020
Cited alongside, same era.
Contemporary abstract algebra
Joseph Gallian · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2021
Cited alongside, same era.
Efficiently Modeling Long Sequences with Structured State Spaces
A. Gu, K. Goel, and C. Re · 2022
Cited alongside, same era.
FLA: A Triton-Based Library for Hardware-Efficient Implementations of Linear Attention Mechanism, January 2024
Songlin Yang and Yu Zhang · 2024
Later among the works it cites.
Advancing Regular Language Reasoning in Linear Recurrent Neural Networks
Ting-Han Fan, Ta-Chung Chi, and Alexander Rudnicky · 2024
Later among the works it cites.
B’MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
L. Zancato, A. Seshadri, Y. Dukler, A. Golatkar, Y. Shen, B. Bowman, M. Trager, A. Achille, and S. Soatto · 2024
Later among the works it cites.
RotRNN: Modelling Long Sequences with Rotations
Kai Biegun, Rares Dolga, Jake Cunningham, and David Barber · 2024
Later among the works it cites.
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale, 2024
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Later among the works it cites.
A framework for few-shot language model evaluation, 07 2024
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2024
Later among the works it cites.
Nnsight and ndif: Democratizing access to foundation model internals
J. Fiotto-Kaufman, A. R. Loftus, E. Todd, J. Brinkmann, K. Pal, D. Troitskii, M. Ripa, A. Belfki, C. Rager, C. Juang, A. Mueller, S. Marks, A. Sen Sharma, F. Lucchetti, N. Prakash, C. Brodley, A. Guha, J. Bell, B. C. Wallace, and D. Bau · 2024
Later among the works it cites.
Gated delta networks: Improving mamba2 with delta rule
S. Yang, J. Kautz, and A. Hatamizadeh · 2025
Closest in time.
RWKV-7 ”Goose” with Expressive Dynamic State Evolution, 2025
Bo Peng, Ruichong Zhang, Daniel Goldstein, Eric Alcaide, Haowen Hou, Janna Lu, William Merrill, Guangyu Song, Kaifeng Tan, Saiteja Utpala, Nathan Wilce, Johan S. Wind, Tianyi Wu, Daniel Wuttke, and Christian Zhou-Zheng · 2025
Closest in time.
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
R. Grazzi, J. Siems, A. Zela, J. Franke, F. Hutter, and M. Pontil · 2025
Closest in time.
Test-time regression: A unifying framework for designing sequence models with associative memory
Ke Alexander Wang, Jiaxin Shi, and Emily B Fox · 2025
Closest in time.
Back to recurrent processing at the crossroad of transformers and state-space models
Matteo Tiezzi, Michele Casoni, Alessandro Betti, Tommaso Guidi, Marco Gori, and Stefano Melacci · 2025
Closest in time.
Scaling up test-time compute with latent reasoning: A recurrent depth approach
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein · 2025
Closest in time.
Implicit Language Models are RNNs: Balancing Parallelization and Expressivity
Mark Schöne, Babak Rahmani, Heiner Kremer, Fabian Falck, Hitesh Ballani, and Jannes Gladrow · 2025
Closest in time.
Fixed-point RNNs: From diagonal to dense in a few iterations
Sajad Movahedi, Felix Sarnthein, Nicola Muca Cirone, and Antonio Orvieto · 2025
Closest in time.
Forgetting transformer: Softmax attention with a forget gate
Zhixuan Lin, Evgenii Nikishin, Xu Owen He, and Aaron Courville · 2025
Closest in time.
Open Thoughts
OpenThoughts Team · 2025
Closest in time.
Quantifying Memory Utilization with Effective State-Size
Rom N Parnichkun, Neehal Tumma, Armin W Thomas, Alessandro Moro, Qi An, Taiji Suzuki, Atsushi Yamashita, Michael Poli, and Stefano Massaroli · 2025
Closest in time.
ParallelFlow: Parallelizing Linear Transformers via Flow Discretization
Nicola Muca Cirone and Cristopher Salvi · 2025
Closest in time.
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
Maximilian Beck, Korbinian Pöppel, Phillip Lippe, and Sepp Hochreiter · 2025
Closest in time.