Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs) have fast inference and scale efficiently on long sequences, but they are difficult to train and hard to scale.
A new approach to linear filtering and prediction problems
R. E. Kalman · 1960
Earlier work this paper cites.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
P. J. Werbos · 1990
Earlier work this paper cites.
Turing computability with neural nets
H. T. Siegelmann and E. D. Sontag · 1991
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2002
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernocký, and S. Khudanpur · 2010
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Quasi-recurrent neural networks
J. Bradbury, S. Merity, C. Xiong, and R. Socher · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Parallelizing linear recurrent neural nets over sequence length
E. Martin and C. Cundy · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Nvidia tensor core programmability, performance & precision
S. Markidis, S. W. Der Chien, E. Laure, I. B. Peng, and J. S. Vetter · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
N. Shazeer · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Cited alongside, same era.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Cited alongside, same era.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
Competition-level code generation with alphacode
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, et al · 2022
Later among the works it cites.
Long range language modeling via gated state spaces
H. Mehta, A. Gupta, A. Cutkosky, and B. Neyshabur · 2022
Later among the works it cites.
Simplified state space layers for sequence modeling
J. T. Smith, A. Warrington, and S. W. Linderman · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hippo: Recurrent memory with optimal polynomial projections
A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré · 2020
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
Transformers are RNNs: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Cited alongside, same era.
Glu variants improve transformer
N. Shazeer · 2020
Cited alongside, same era.
Long range arena: A benchmark for efficient transformers
Y. Tay, M. Dehghani, S. Abnar, Y. Shen, D. Bahri, P. Pham, J. Rao, L. Yang, S. Ruder, and D. Metzler · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu · 2020
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, et al · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team Google · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
A. Gu and T. Dao · 2023
Later among the works it cites.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Later among the works it cites.
Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings
N. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, et al · 2023
Later among the works it cites.
Gateloop: Fully data-controlled linear recurrence for sequence modeling
T. Katsch · 2023
Later among the works it cites.
Rwkv: Reinventing RNNs for the transformer era
B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, H. Cao, X. Cheng, M. Chung, M. Grella, K. K. GV, et al · 2023
Later among the works it cites.
Hyena hierarchy: Towards larger convolutional language models
M. Poli, S. Massaroli, E. Nguyen, D. Y. Fu, T. Dao, S. Baccus, Y. Bengio, S. Ermon, and C. Ré · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Y. Sun, L. Dong, S. Huang, S. Ma, Y. Xia, J. Xue, J. Wang, and F. Wei · 2023
Later among the works it cites.
LLama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Repeat after me: Transformers are better than state space models at copying
S. Jelassi, D. Brandfonbrener, S. M. Kakade, and E. Malach · 2024
Closest in time.
The impact of positional encoding on length generalization in transformers
A. Kazemnejad, I. Padhi, K. Natesan Ramamurthy, P. Das, and S. Reddy · 2024
Closest in time.
Mambabyte: Token-free selective state space model
J. Wang, T. Gangavarapu, J. N. Yan, and A. M. Rush · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang · 2024
Closest in time.