Fetching the paper…
Reading the bibliography…
It is notoriously difficult to train Transformers on small datasets; typically, large pre-trained models are instead used as the starting point.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Bai, S., Kolter, J. Z., and Koltun, V · 2018
Earlier work this paper cites.
On the relationship between self-attention and convolutional layers
Cordonnier, J.-B., Loukas, A., and Jaggi, M · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Improving transformer optimization through better initialization
Huang, X. S., Perez, F., Ba, J., and Volkovs, M · 2020
Earlier work this paper cites.
Coatnet: Marrying convolution and attention for all data sizes
Dai, Z., Liu, H., Le, Q. V., and Tan, M · 2021
Earlier work this paper cites.
Convit: Improving vision transformers with soft convolutional inductive biases
d’Ascoli, S., Touvron, H., Leavitt, M. L., Morcos, A. S., Biroli, G., and Sagun, L · 2021
Earlier work this paper cites.
Escaping the big data paradigm with compact transformers
Hassani, A., Walton, S., Shah, N., Abuduweili, A., Li, J., and Shi, H · 2021
Cited alongside, same era.
Vision transformer for small-size datasets
Lee, S. H., Lee, S., and Song, B. C · 2021
Cited alongside, same era.
Efficient training of visual transformers with small datasets
Liu, Y., Sangineto, E., Bi, W., Sebe, N., Lepri, B., and Nadai, M · 2021
Cited alongside, same era.
Resnet strikes back: An improved training procedure in timm
Wightman, R., Touvron, H., and Jégou, H · 2021
Cited alongside, same era.
Cvt: Introducing convolutions to vision transformers
Wu, H., Xiao, B., Codella, N., Liu, M., Dai, X., Yuan, L., and Zhang, L · 2021
Cited alongside, same era.
Training vision transformers with only 2040 images
Cao, Y.-H., Yu, H., and Wu, J · 2022
Later among the works it cites.
How to train vision transformer on small-scale datasets?
Gani, H., Naseer, M., and Yaqub, M · 2022
Later among the works it cites.
Trockman, A. and Kolter, J. Z · 2022
Later among the works it cites.
Understanding the covariance structure of convolutional filters
Trockman, A., Willmott, D., and Kolter, J. Z · 2022
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Incorporating convolution designs into visual transformers
Yuan, K., Guo, S., Liu, Z., Zhou, A., Yu, F., and Wu, W · 2021
Cited alongside, same era.
Zero initialization: Initializing residual networks with only zeros and ones
Zhao, J., Schäfer, F., and Anandkumar, A · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H
Cited in the paper.
Going deeper with image transformers
Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., and Jégou, H
Cited in the paper.
Zhang, Y., Backurs, A., Bubeck, S., Eldan, R., Gunasekar, S., and Wagner, T · 2022
Later among the works it cites.
Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation
He, B., Martens, J., Zhang, G., Botev, A., Brock, A., Smith, S. L., and Teh, Y. W · 2023
Closest in time.