Fetching the paper…
Reading the bibliography…
Recently, foundation models based on Vision Transformers (ViTs) have become widely available.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Dodge, J.; Ilharco, G.; Schwartz, R.; Farhadi, A.; Hajishirzi, H.; and Smith, N. 2020 · 2002
Earlier work this paper cites.
Automated Flower Classification over a Large Number of Classes
Nilsback, M.-E.; and Zisserman, A. 2008 · 2008
Earlier work this paper cites.
CIFAR-100 (Canadian Institute for Advanced Research)
Krizhevsky, A.; Nair, V.; and Hinton, G. 2009 · 2009
Earlier work this paper cites.
Food-101 – Mining Discriminative Components with Random Forests
Bossard, L.; Guillaumin, M.; and Van Gool, L. 2014 · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; Berg, A. C.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Training Deep Nets with Sublinear Memory Cost
Chen, T.; Xu, B.; Zhang, C.; and Guestrin, C. 2016 · 2016
Earlier work this paper cites.
Memory-efficient backpropagation through time
Gruslys, A.; Munos, R.; Danihelka, I.; Lanctot, M.; and Graves, A. 2016 · 2016
Earlier work this paper cites.
Automatic differentiation in PyTorch
Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2017 · 2017
Earlier work this paper cites.
Mixed Precision Training
Micikevicius, P.; Narang, S.; Alben, J.; Diamos, G.; Elsen, E.; Garcia, D.; Ginsburg, B.; Houston, M.; Kuchaiev, O.; Venkatesh, G.; and Wu, H. 2018 · 2018
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
PyTorch Image Models
Wightman, R. 2019 · 2019
Earlier work this paper cites.
TinyTL: Reduce Memory, Not Parameters for Efficient On-Device Learning
Cai, H.; Gan, C.; Zhu, L.; and Han, S. 2020 · 2020
Earlier work this paper cites.
SPINN: synergistic progressive inference of neural networks over device and cloud
Laskaridis, S.; Venieris, S. I.; Almeida, M.; Leontiadis, I.; and Lane, N. D. 2020 · 2020
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
A Mathematical Framework for Transformer Circuits
Elhage, N.; Nanda, N.; Olsson, C.; Henighan, T.; Joseph, N.; Mann, B.; Askell, A.; Bai, Y.; Chen, A.; Conerly, T.; DasSarma, N.; Drain, D.; Ganguli, D.; Hatfield-Dodds, Z.; Hernandez, D.; Jones, A.; Kernion, J.; Lovitt, L.; Ndousse, K.; Amodei, D.; Brown, T.; Clark, J.; Kaplan, J.; McCandlish, S.; and Olah, C. 2021 · 2021
Cited alongside, same era.
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Cited alongside, same era.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer
Mehta, S.; and Rastegari, M. 2022 · 2022
Later among the works it cites.
Adavit: Adaptive vision transformers for efficient image recognition
Meng, L.; Li, H.; Chen, B.-C.; Lan, S.; Wu, Z.; Jiang, Y.-G.; and Lim, S.-N. 2022 · 2022
Later among the works it cites.
A-vit: Adaptive tokens for efficient vision transformer
Yin, H.; Vahdat, A.; Alvarez, J. M.; Mallya, A.; Kautz, J.; and Molchanov, P. 2022 · 2022
Later among the works it cites.
Are All Layers Created Equal?
Zhang, C.; Bengio, S.; and Singer, Y. 2022 · 2022
Later among the works it cites.
Token Merging: Your ViT but Faster
Bolya, D.; Fu, C.-Y.; Dai, X.; Zhang, P.; Feichtenhofer, C.; and Hoffman, J. 2023 · 2023
Later among the works it cites.
Visual Instruction Tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Do Vision Transformers See Like Convolutional Neural Networks?
Raghu, M.; Unterthiner, T.; Kornblith, S.; Zhang, C.; and Dosovitskiy, A. 2021 · 2021
Cited alongside, same era.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Rao, Y.; Zhao, W.; Liu, B.; Lu, J.; Zhou, J.; and Hsieh, C.-J. 2021 · 2021
Cited alongside, same era.
Training data-efficient image transformers with distillation through attention
Touvron, H.; Cord, M.; Douze, M.; Massa, F.; Sablayrolles, A.; and Jegou, H. 2021 · 2021
Cited alongside, same era.
Efficientvit: Enhanced linear attention for high-resolution low-computation visual recognition
Cai, H.; Gan, C.; and Han, S. 2022 · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; yelong shen; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Cited alongside, same era.
EViT: Expediting Vision Transformers via Token Reorganizations
Liang, Y.; GE, C.; Tong, Z.; Song, Y.; Wang, J.; and Xie, P. 2022 · 2022
Cited alongside, same era.
FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer
Lin, Y.; Zhang, T.; Sun, P.; Li, Z.; and Zhou, S. 2022 · 2022
Cited alongside, same era.
Weight subcloning: direct initialization of transformers using larger pretrained ones
Samragh, M.; Farajtabar, M.; Mehta, S.; Vemulapalli, R.; Faghri, F.; Naik, D.; Tuzel, O.; and Rastegari, M. 2023 · 2023
Later among the works it cites.
Adaptive Computation Modules: Granular Conditional Computation For Efficient Inference
Wójcik, B.; Devoto, A.; Pustelnik, K.; Minervini, P.; and Scardapane, S. 2023 · 2023
Later among the works it cites.
Goal-oriented and semantic communication in 6G AI-native networks: The 6G-GOALS approach
Calvanese Strinati, E.; Di Lorenzo, P.; Sciancalepore, V.; Aijaz, A.; Kountouris, M.; Gündüz, D.; Popovski, P.; Sana, M.; Stavrou, P. A.; and et al. 2024 · 2024
Closest in time.
The Unreasonable Ineffectiveness of the Deeper Layers
Gromov, A.; Tirumala, K.; Shapourian, H.; Glorioso, P.; and Roberts, D. A. 2024 · 2024
Closest in time.
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
Jain, G.; Hegde, N.; Kusupati, A.; Nagrani, A.; Buch, S.; Jain, P.; Arnab, A.; and Paul, S. 2024 · 2024
Closest in time.
Block Selective Reprogramming for On-device Training of Vision Transformers
Sarkar, S.; Kundu, S.; Zheng, K.; and Beerel, P. A. 2024 · 2024
Closest in time.
Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey
Xin, Y.; Luo, S.; Zhou, H.; Du, J.; Liu, X.; Fan, Y.; Li, Q.; and Du, Y. 2024 · 2024
Closest in time.