Fetching the paper…
Reading the bibliography…
Recurrent Neural Networks (RNNs) offer fast inference on long sequences but are hard to optimize and slow to train.
Dynamical systems of continuous spectra
B. O. Koopman and J. v. Neumann · 1932
Earlier work this paper cites.
A logical calculus of the ideas immanent in nervous activity
W. S. McCulloch and W. Pitts · 1943
Earlier work this paper cites.
Statistical ensembles of complex, quaternion, and real matrices
J. Ginibre · 1965
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
J. J. Hopfield · 1982
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1985
Earlier work this paper cites.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronales netzen
S. Hochreiter · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
The dynamic universality of sigmoidal neural networks
J. Kilian and H. T. Siegelmann · 1996
Earlier work this paper cites.
Linear algebra done right
S. Axler · 1997
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
The theory of probability
H. Jeffreys · 1998
Earlier work this paper cites.
Computational methods for inverse problems
C. R. Vogel · 2002
Earlier work this paper cites.
The jordan canonical form of a rational random matrix
Z. Zhinan · 2002
Earlier work this paper cites.
Backpropagation-decorrelation: online recurrent learning with o (n) complexity
J. J. Steil · 2004
Earlier work this paper cites.
Jordan canonical form: theory and practice
S. H. Weintraub · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Dynamic mode decomposition of numerical and experimental data
P. J. Schmid · 2010
Earlier work this paper cites.
Neural networks and analog computation: beyond the Turing limit
H. T. Siegelmann · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Q. V. Le, N. Jaitly, and G. E. Hinton · 2015
Earlier work this paper cites.
A data–driven approximation of the koopman operator: Extending dynamic mode decomposition
M. O. Williams, I. G. Kevrekidis, and C. W. Rowley · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
M. Arjovsky, A. Shah, and Y. Bengio · 2016
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
Neural machine translation in linear time
N. Kalchbrenner, L. Espeholt, K. Simonyan, A. v. d. Oord, A. Graves, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Dynamic mode decomposition: data-driven modeling of complex systems
J. N. Kutz, S. L. Brunton, B. W. Brunton, and J. L. Proctor · 2016
Cited alongside, same era.
Global stability analysis using the eigenfunctions of the koopman operator
A. Mauroy and I. Mezić · 2016
Cited alongside, same era.
Legendre memory units: Continuous-time representation in recurrent neural networks
A. Voelker, I. Kajić, and C. Eliasmith · 2019
Later among the works it cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, S. Shyam, G. Sastry, A. Askell, et al · 2020
Later among the works it cites.
Batch normalization biases residual blocks towards the identity function in deep networks
S. De and S. Smith · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, N. Houlsby, S. Gelly, X. Zhang, and J. Uszkoreit · 2020
Later among the works it cites.
Hippo: Recurrent memory with optimal polynomial projections
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Nallapati, B. Zhou, C. Gulcehre, B. Xiang, et al · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Koopman operator based observer synthesis for control-affine nonlinear systems
A. Surana · 2016
Cited alongside, same era.
Full-capacity unitary recurrent neural networks
S. Wisdom, T. Powers, J. Hershey, J. Le Roux, and L. Atlas · 2016
Cited alongside, same era.
Language modeling with gated convolutional networks
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier · 2017
Cited alongside, same era.
Learning unitary operators with help from u (n)
S. L. Hyland and G. Rätsch · 2017
Cited alongside, same era.
Tunable efficient unitary neural networks (eunn) and their application to rnns
L. Jing, Y. Shen, T. Dubcek, J. Peurifoy, S. Skirlo, Y. LeCun, M. Tegmark, and M. Soljačić · 2017
Cited alongside, same era.
Haiku: Sonnet for JAX, 2020
T. Hennigan, T. Cai, T. Norman, and I. Babuschkin · 2020
Later among the works it cites.
Koopman model predictive control of nonlinear dynamical systems
M. Korda and I. Mezić · 2020
Later among the works it cites.
Koopman operator in systems and control
A. Mauroy, Y. Susuki, and I. Mezić · 2020
Later among the works it cites.
Long range arena: A benchmark for efficient transformers
Y. Tay, M. Dehghani, S. Abnar, Y. Shen, D. Bahri, P. Pham, J. Rao, L. Yang, S. Ruder, and D. Metzler · 2020
Later among the works it cites.
Turing completeness of bounded-precision recurrent neural networks
S. Chung and H. Siegelmann · 2021
Later among the works it cites.
Lipschitz recurrent neural networks
N. B. Erichson, O. Azencot, A. Queiruga, L. Hodgkinson, and M. W. Mahoney · 2021
Later among the works it cites.
Liquid time-constant networks
R. Hasani, M. Lechner, A. Amini, D. Rus, and R. Grosu · 2021
Later among the works it cites.
Highly accurate protein structure prediction with alphafold
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, et al · 2021
Later among the works it cites.
Data-driven discovery of koopman eigenfunctions for control
E. Kaiser, J. N. Kutz, and S. L. Brunton · 2021
Later among the works it cites.
Fnet: Mixing tokens with fourier transforms
J. Lee-Thorp, J. Ainslie, I. Eckstein, and S. Ontanon · 2021
Later among the works it cites.
Novel machine learning approaches revolutionize protein knowledge
N. Bordin, C. Dallago, M. Heinzinger, S. Kim, M. Littmann, C. Rauer, M. Steinegger, B. Rost, and C. Orengo · 2022
Later among the works it cites.
It’s raw! audio generation with state-space models
K. Goel, A. Gu, C. Donahue, and C. Ré · 2022
Later among the works it cites.
Liquid structural state-space models
R. Hasani, M. Lechner, T.-H. Wang, M. Chahine, A. Amini, and D. Rus · 2022
Later among the works it cites.
Long movie clip classification with state-space video models
M. M. Islam and G. Bertasius · 2022
Later among the works it cites.
Learning dynamical systems via koopman operator regression in reproducing kernel hilbert spaces
V. R. Kostic, P. Novelli, A. Maurer, C. Ciliberto, L. Rosasco, and massimiliano pontil · 2022
Later among the works it cites.
Mega: moving average equipped gated attention
X. Ma, C. Zhou, X. Kong, J. He, L. Gui, G. Neubig, J. May, and L. Zettlemoyer · 2022
Later among the works it cites.
Long range language modeling via gated state spaces
H. Mehta, A. Gupta, A. Cutkosky, and B. Neyshabur · 2022
Later among the works it cites.
S4nd: Modeling images and videos as multidimensional signals with state spaces
E. Nguyen, K. Goel, A. Gu, G. Downs, P. Shah, T. Dao, S. Baccus, and C. Ré · 2022
Later among the works it cites.
Simplified state space layers for sequence modeling
J. T. Smith, A. Warrington, and S. W. Linderman · 2022
Later among the works it cites.
The effects of nonlinearity on approximation capacity of recurrent neural networks, 2022
S. Wang, Z. Li, and Q. Li · 2022
Later among the works it cites.