Fetching the paper…
Reading the bibliography…
The self-attention mechanism prevails in modern machine learning.
Generating sequences with recurrent neural networks
Graves, A · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Martins, A. and Astudillo, R · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Cited alongside, same era.
ReZero is all you need: Fast convergence at large depth
Bachlechner, T., Majumder, B. P., Mao, H., Cottrell, G., and McAuley, J · 2021
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong, Y., Cordonnier, J.-B., and Loukas, A · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Cited alongside, same era.
What they do when in doubt: a study of inductive biases in seq2seq learners
Kharitonov, E. and Chaabouni, R · 2021
Cited alongside, same era.
slimIPL: Language-model-free iterative pseudo-labeling
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Noci, L., Anagnostidis, S., Biggio, L., Orvieto, A., Singh, S. P., and Lucchi, A · 2022
Later among the works it cites.
An explanation of in-context learning as implicit Bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2022
Later among the works it cites.
Birth of a transformer: A memory viewpoint
Bietti, A., Cabannes, V., Bouchacourt, D., Jegou, H., and Bottou, L · 2023
Later among the works it cites.
The matrix reference manual
Brookes, M · 2023
Later among the works it cites.
Analyzing feed-forward blocks in transformers through the lens of attention map
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Likhomanenko, T., Xu, Q., Kahn, J., Synnaeve, G., and Collobert, R · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Cited alongside, same era.
Vision transformers provably learn spatial structure
Jelassi, S., Sander, M., and Li, Y · 2022
Cited alongside, same era.
Grokking of hierarchical structure in vanilla transformers
Murty, S., Sharma, P., Andreas, J., and Manning, C
Cited in the paper.
Characterizing intrinsic compositionality in transformers with tree projections
Murty, S., Sharma, P., Andreas, J., and Manning, C. D
Cited in the paper.
Spike no more: Stabilizing the pre-training of large language models
Takase, S., Kiyono, S., Kobayashi, S., and Suzuki, J
Cited in the paper.
Transformers as support vector machines
Tarzanagh, D. A., Li, Y., Thrampoulidis, C., and Oymak, S
Cited in the paper.
Kobayashi, G., Kuribayashi, T., Yokoi, S., and Inui, K · 2023
Later among the works it cites.
B2T connection: Serving stability and performance in deep transformers
Takase, S., Kiyono, S., Kobayashi, S., and Suzuki, J · 2023
Later among the works it cites.
Stabilizing transformer training by preventing attention entropy collapse
Zhai, S., Likhomanenko, T., Littwin, E., Busbridge, D., Ramapuram, J., Zhang, Y., Gu, J., and Susskind, J. M · 2023
Later among the works it cites.