Fetching the paper…
Reading the bibliography…
Recent architectural developments have enabled recurrent neural networks (RNNs) to reach and even surpass the performance of Transformers on certain sequence modeling tasks.
Die lernmatrix
Steinbuch, K · 1961
Earlier work this paper cites.
Non-holographic associative memory
Willshaw, D. J., Buneman, O. P., and Longuet-Higgins, H. C · 1969
Earlier work this paper cites.
Correlation matrix memories
Kohonen, T · 1972
Earlier work this paper cites.
Fading memory and the problem of approximating nonlinear operators with Volterra series
Boyd, S. and Chua, L · 1985
Earlier work this paper cites.
Learning to control fast-weight memories: an alternative to dynamic recurrent networks
Schmidhuber, J · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Theoretical neuroscience: computational and mathematical modeling of neural systems
Dayan, P. and Abbott, L. F · 2001
Earlier work this paper cites.
Matplotlib: A 2D graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Neuronal arithmetic
Silver, R. A · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., and others · 2011
Earlier work this paper cites.
On the properties of neural machine translation: encoder-decoder approaches
Cho, K., van Merrienboer, B., Bahdanau, D., and Bengio, Y · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Using fast weights to attend to the recent past
Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Universal discrete-time reservoir computers with stochastic inputs and linear readouts using non-homogeneous state-affine systems
Grigoryeva, L. and Ortega, J.-P · 2018
Cited alongside, same era.
Differentiable plasticity: training plastic neural networks with backpropagation
Miconi, T., Clune, J., and Stanley, K. O · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Transformer dissection: a unified understanding of transformer’s attention via the lens of kernel
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., and others · 2020
Cited alongside, same era.
Array programming with NumPy
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S · 2023
Closest in time.
A practical survey on faster and lighter transformers
Fournier, Q., Caron, G. M., and Aloise, D · 2023
Closest in time.
Hungry Hungry Hippos: Towards Language Modeling with State Space Models
Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Ré, C · 2023
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
Gu, A. and Dao, T · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Harris, C. R., Millman, K. J., Walt, S. J. v. d., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., Kerkwijk, M. H. v., Brett, M., Haldane, A., Río, J. F. d., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Cited alongside, same era.
Transformers are RNNs: fast autoregressive Transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2020
Cited alongside, same era.
Rethinking attention with Performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., Belanger, D., Colwell, L., and Weller, A · 2021
Cited alongside, same era.
Random feature attention
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L · 2021
Cited alongside, same era.
Linear Transformers are secretly fast weight programmers
Schlag, I., Irie, K., and Schmidhuber, J · 2021
Cited alongside, same era.
Efficient attention: attention with linear complexities
Shen, Z., Zhang, M., Zhao, H., Yi, S., and Li, H · 2021
Cited alongside, same era.
Flax: A neural network library and ecosystem for JAX, 2023
Heek, J., Levskaya, A., Oliver, A., Ritter, M., Rondepierre, B., Steiner, A., and Zee, M. v · 2023
Closest in time.
Toward a formal theory for computing machines made out of whatever physics offers
Jaeger, H., Noheda, B., and Van Der Wiel, W. G · 2023
Closest in time.
Mahankali, A., Hashimoto, T. B., and Ma, T · 2023
Closest in time.
Expand-and-cluster: exact parameter recovery of neural networks
Martinelli, F., Simsek, B., Brea, J., and Gerstner, W · 2023
Closest in time.
On the universality of linear recurrences followed by nonlinear projections
Orvieto, A., De, S., Gulcehre, C., Pascanu, R., and Smith, S. L · 2023
Closest in time.
RWKV: Reinventing RNNs for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., He, X., Hou, H., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Mantri, K. S. I., Mom, F., Saito, A., Tang, X., Wang, B., Wind, J. S., Wozniak, S., Zhang, R., Zhang, Z., Zhao, Q., Zhou, P., Zhu, J., and Zhu, R.-J · 2023
Closest in time.
Simplified state space layers for sequence modeling
Smith, J. T., Warrington, A., and Linderman, S. W · 2023
Closest in time.
Retentive network: A successor to transformer for large language models, 2023
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F · 2023
Closest in time.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Closest in time.
Gated linear attention Transformers with hardware-efficient training, 2023
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2023
Closest in time.
Trained transformers learn linear models in-context
Zhang, R., Frei, S., and Bartlett, P. L · 2023
Closest in time.