Fetching the paper…
Reading the bibliography…
The Transformer architecture has revolutionized deep learning through its Self-Attention mechanism, which effectively captures contextual information.
Fast Transformer Decoding: One Write-Head is All You Need
Shazeer, N. 2019 · 1911
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. 2009 · 2009
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2010
Earlier work this paper cites.
Food-101 – Mining Discriminative Components with Random Forests
Bossard, L.; Guillaumin, M.; and Van Gool, L. 2014 · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; Berg, A. C.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Tiny ImageNet
mnmoustafa, M. A. 2017 · 2017
Earlier work this paper cites.
CINIC-10 is not ImageNet or CIFAR-10
Darlow, L. N.; Crowley, E. J.; Antoniou, A.; and Storkey, A. J. 2018 · 2018
Cited alongside, same era.
PyTorch Image Models
Wightman, R. 2019 · 2019
Cited alongside, same era.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Cited alongside, same era.
CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
Dong, X.; Bao, J.; Chen, D.; Zhang, W.; Yu, N.; Yuan, L.; Chen, D.; and Guo, B. 2022 · 2022
Cited alongside, same era.
Training Compute-Optimal Large Language Models
Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; de Las Casas, D.; Hendricks, L. A.; Welbl, J.; Clark, A.; Hennigan, T.; Noland, E.; Millican, K.; van den Driessche, G.; Damoc, B.; Guy, A.; Osindero, S.; Simonyan, K.; Elsen, E.; Rae, J. W.; Vinyals, O.; and Sifre, L. 2022 · 2022
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
Javadi, F.; Ahmed, W.; Hajimolahoseini, H.; Ataiefard, F.; Hassanpour, M.; Asani, S.; Wen, A.; Awad, O. M.; Liu, K.; and Liu, Y. 2023 · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023 · 2023
Later among the works it cites.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2023 · 2023
Later among the works it cites.
Optimised Grouped-Query Attention Mechanism for Transformers
Chen, Y.; Zhang, C.; Gao, X.; Mullins, R. D.; Constantinides, G. A.; and Zhao, Y. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
Steiner, A.; Kolesnikov, A.; Zhai, X.; Wightman, R.; Uszkoreit, J.; and Beyer, L. 2022 · 2022
Cited alongside, same era.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Ainslie, J.; Lee-Thorp, J.; de Jong, M.; Zemlyanskiy, Y.; Lebrón, F.; and Sanghai, S. 2023 · 2023
Cited alongside, same era.
Chinnakonduru, S. S.; and Mohapatra, A. 2024 · 2024
Closest in time.
QCQA: Quality and Capacity-aware grouped Query Attention
Joshi, V.; Laddha, P.; Sinha, S.; Omer, O. J.; and Subramoney, S. 2024 · 2024
Closest in time.