Fetching the paper…
Reading the bibliography…
Transformers have achieved state-of-the-art performance in language modeling tasks.
On early stopping in gradient descent learning
Yao, Y., Rosasco, L., and Caponnetto, A · 2007
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Bird, S., Klein, E., and Loper, E · 2009
Earlier work this paper cites.
Speech and language processing : an introduction to natural language processing, computational linguistics, and speech recognition
Jurafsky, D. and Martin, J. H · 2009
Earlier work this paper cites.
Perturbation theory for linear operators , volume 132
Kato, T · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Arora, S., Cohen, N., Hu, W., and Luo, Y · 2019
Earlier work this paper cites.
Orthogonality constrained multi-head attention for keyword spotting
Lee, M., Lee, J., Jang, H. J., Kim, B., Chang, W., and Hwang, K · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S., Tsipras, D., Liang, P. S., and Valiant, G · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
Jelassi, S., Sander, M., and Li, Y · 2022
Later among the works it cites.
Sinkformers: Transformers with doubly stochastic attention
Sander, M. E., Ablin, P., Blondel, M., and Peyré, G · 2022
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Cited alongside, same era.
The loss landscape of deep linear neural networks: a second-order analysis
Achour, E. M., Malgouyres, F., and Gerchinovitz, S · 2021
Cited alongside, same era.
Convergence of gradient descent for learning linear neural networks
Nguegnang, G. M., Rauhut, H., and Terstiege, U · 2021
Cited alongside, same era.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Pesme, S., Pillaud-Vivien, L., and Flammarion, N · 2021
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Cited alongside, same era.
On orthogonality constraints for transformers
Zhang, A., Chan, A., Tay, Y., Fu, J., Wang, S., Zhang, S., Shao, H., Yao, S., and Lee, R. K.-W · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Later among the works it cites.
A practical survey on faster and lighter transformers
Fournier, Q., Caron, G. M., and Aloise, D · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S · 2023
Later among the works it cites.
Transformers as algorithms: Generalization and stability in in-context learning
Li, Y., Ildiz, M. E., Papailiopoulos, D., and Oymak, S · 2023
Later among the works it cites.
Mahankali, A., Hashimoto, T. B., and Ma, T · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Zhang, R., Frei, S., and Bartlett, P. L · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.