Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Shifting inductive bias with success-story algorithm, adaptive levin search, and incremental self-improvement
Schmidhuber, J., Zhao, J., and Wiering, M · 1997
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H · 2016
Earlier work this paper cites.
Fast learning requires good memory: A time-space lower bound for parity learning
Raz, R · 2016
Earlier work this paper cites.
Time-space hardness of learning sparse parities
Kol, G., Raz, R., and Tal, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Original
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Theoretical limitations of self-attention in neural sequence models
Hahn, M · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Original
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Original
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Earlier work this paper cites.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2021
Earlier work this paper cites.