Fetching the paper…
Reading the bibliography…
The success of large-scale language models like GPT can be attributed to their ability to efficiently predict the next token in a sequence.
Elbayad, M., Zeghidour, N., and Usunier, N. (2019) · 1910
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Graves, A. (2016) · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Universal transformers
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L. (2019) · 2019
Cited alongside, same era.
Openwebtext corpus
Gokaslan, A. and Cohen, V. (2019) · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Later among the works it cites.
Lessons on parameter sharing across layers in transformers
Takase, S. and Kiyono, S. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…