Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs) notoriously struggle to learn long-term memories, primarily due to vanishing and exploding gradients.
Fading memory and the problem of approximating nonlinear operators with Volterra series
S. Boyd and L. Chua · 1985
Earlier work this paper cites.
Sequential thought processes in PDP models
David E Rumelhart, Paul Smolensky, James L McClelland, and G Hinton · 1986
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
Sepp Hochreiter · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
Salah El Hihi and Yoshua Bengio · 1995
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The columnar organization of the neocortex
V B Mountcastle · 1997
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, and Jurgen Schmidhuber · 2001
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Learning recurrent neural networks with hessian-free optimization
James Martens and Ilya Sutskever · 2011
Earlier work this paper cites.
Statistical language models based on neural networks
Tomas Mikolov · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
On the properties of neural machine translation: encoder-decoder approaches
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Stochastic processes and applications , volume 60 of
Grigorios A. Pavliotis · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Batch normalization: accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc V. Le, Navdeep Jaitly, and Geoffrey E. Hinton · 2015
Earlier work this paper cites.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Adam: a method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Recurrent batch normalization
Tim Cooijmans, Nicolas Ballas, César Laurent, Çağlar Gülçehre, and Aaron Courville · 2017
Cited alongside, same era.
Fast-slow recurrent neural networks
Asier Mujika, Florian Meier, and Angelika Steger · 2017
Cited alongside, same era.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Cited alongside, same era.
FlashAttention: fast and memory-efficient exact attention with IO-awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Later among the works it cites.
Switch Transformers: scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Later among the works it cites.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2022
Later among the works it cites.
Diagonal state spaces are as effective as structured states spaces
Ankit Gupta, Albert Gu, and Jonathan Berant · 2022
Later among the works it cites.
Approximation and optimization theory for linear continuous-time recurrent neural networks
Zhong Li, Jiequn Han, Weinan E, and Qianxiao Li · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the computation of complex-valued gradients with application to statistically optimum beamforming
Christoph Boeddeker, Patrick Hanebrink, Lukas Drude, Jahn Heymann, and Reinhold Haeb-Umbach · 2017
Cited alongside, same era.
Can recurrent neural networks warp time?
Corentin Tallec and Yann Ollivier · 2018
Cited alongside, same era.
Orthogonal recurrent neural networks with scaled Cayley transform
Kyle Helfrich, Devin Willmott, and Qiang Ye · 2018
Cited alongside, same era.
Dynamical isometry and a mean field theory of RNNs: gating enables signal propagation in recurrent neural networks
Minmin Chen, Jeffrey Pennington, and Samuel S. Schoenholz · 2018
Cited alongside, same era.
Gradient descent learns linear dynamical systems
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Yurii Nesterov and others · 2018
Cited alongside, same era.
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2022
Later among the works it cites.
A practical survey on faster and lighter transformers
Quentin Fournier, Gaétan Marceau Caron, and Daniel Aloise · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
Jimmy T.H. Smith, Andrew Warrington, and Scott W. Linderman · 2023
Later among the works it cites.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Çağlar Gülçehre, Razvan Pascanu, and Soham De · 2023
Later among the works it cites.
Mamba: linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Neural wave machines: learning spatiotemporally structured representations with locally coupled oscillatory recurrent neural networks
T Anderson Keller and Max Welling · 2023
Later among the works it cites.
Persistent learning signals and working memory without continuous attractors
Il Memming Park, Ábel Ságodi, and Piotr Aleksander Sokół · 2023
Later among the works it cites.
A spectral condition for feature learning
Greg Yang, James B Simon, and Jeremy Bernstein · 2023
Later among the works it cites.
Toward understanding why Adam converges faster than SGD for transformers
Yan Pan and Yuanzhi Li · 2023
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2023
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2023
Later among the works it cites.
The era of 1-bit LLMs: all Large Language Models are in 1.58 bits
Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei · 2024
Closest in time.
Griffin: mixing gated linear recurrences with local attention for efficient language models
Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, Arnaud Doucet, David Budden, Yee Whye Teh, Razvan Pascanu, Nando De Freitas, and Caglar Gulcehre · 2024
Closest in time.
Inverse approximation theory for nonlinear recurrent neural networks
Shida Wang, Zhong Li, and Qianxiao Li · 2024
Closest in time.
Why do learning rates transfer? Reconciling optimization and scaling limits for deep learning
Lorenzo Noci, Alexandru Meterez, Thomas Hofmann, and Antonio Orvieto · 2024
Closest in time.
Universality of linear recurrences followed by non-linear projections: finite-width guarantees and benefits of complex eigenvalues
Antonio Orvieto, Soham De, Çağlar Gülçehre, Razvan Pascanu, and Samuel L. Smith · 2024
Closest in time.
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal · 2024
Closest in time.