Fetching the paper…
Reading the bibliography…
Transformer-based models have recently become wildly successful across a diverse set of domains.
On the kronecker product
Schacke, K · 2004
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Roy, O. and Vetterli, M · 2007
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, 2015
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2018
Earlier work this paper cites.
Deeper insights into graph convolutional networks for semi-supervised learning
Li, Q., Han, Z., and Wu, X.-M · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization, 2018
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2018
Earlier work this paper cites.
Randaugment: Practical automated data augmentation with a reduced search space, 2019
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C. and Toutanova, L. K · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Molecular transformer: A model for uncertainty-calibrated chemical reaction prediction
Schwaller, P., Laino, T., Gaudin, T., Bolgar, P., Hunter, C. A., Bekas, C., and Lee, A. A · 2019
Earlier work this paper cites.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 2019
Earlier work this paper cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features, 2019
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems, 2020
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong, Y., Cordonnier, J.-B., and Loukas, A · 2021
Touvron, H., Cord, M., and Jegou, H · 2022
Later among the works it cites.
Anti-oversmoothing in deep vision transformers via the fourier domain analysis: From theory to practice
Wang, P., Zheng, W., Chen, T., and Wang, Z · 2022
Later among the works it cites.
Addressing token uniformity in transformers via singular value transformation
Yan, H., Gui, L., Li, W., and He, Y · 2022
Later among the works it cites.
Linear attention is (maybe) all you need (to understand transformer optimization)
Ahn, K., Cheng, X., Song, M., Yun, C., Sra, S., and Jadbabaie, A · 2023
Later among the works it cites.
Centered self-attention layers
Ali, A., Galanti, T., and Wolf, L · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Vision transformers with patch diversification
Gong, C., Wang, D., Li, M., Chandra, V., and Liu, Q · 2021
Cited alongside, same era.
Transformer in transformer
Han, K., Xiao, A., Wu, E., Guo, J., Xu, C., and Wang, Y · 2021
Cited alongside, same era.
Token labeling: Training a 85.5% top-1 accuracy vision transformer with 56m parameters on imagenet
Jiang, Z., Hou, Q., Yuan, L., Zhou, D., Jin, X., Wang, A., and Feng, J · 2021
Cited alongside, same era.
Do vision transformers see like convolutional neural networks?
Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., and Dosovitskiy, A · 2021
Cited alongside, same era.
Augmented shortcuts for vision transformers
Tang, Y., Han, K., Xu, C., Xiao, A., Deng, Y., Xu, C., and Wang, Y · 2021
Cited alongside, same era.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.-H., Tay, F. E., Feng, J., and Yan, S · 2021
Cited alongside, same era.
Improving vision transformers by revisiting high-frequency components
Bai, J., Yuan, L., Xia, S.-T., Yan, S., Li, Z., and Liu, W · 2022
Cited alongside, same era.
Choi, J., Wi, H., Kim, J., Shin, Y., Lee, K., Trask, N., and Park, N · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Later among the works it cites.
Cramming: Training a language model on a single gpu in one day
Geiping, J. and Goldstein, T · 2023
Later among the works it cites.
Contranorm: A contrastive learning perspective on oversmoothing and beyond
Guo, X., Wang, Y., Du, T., and Wang, Y · 2023
Later among the works it cites.
One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention
Mahankali, A. V., Hashimoto, T., and Ma, T · 2023
Later among the works it cites.
Matrix analysis and applied linear algebra
Meyer, C. D. and Stewart, I · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Later among the works it cites.
Stabilizing transformer training by preventing attention entropy collapse
Zhai, S., Likhomanenko, T., Littwin, E., Busbridge, D., Ramapuram, J., Zhang, Y., Gu, J., and Susskind, J. M · 2023
Later among the works it cites.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S · 2024
Closest in time.
Trained transformers learn linear models in-context
Zhang, R., Frei, S., and Bartlett, P. L · 2024
Closest in time.