Fetching the paper…
Reading the bibliography…
Convolutional neural networks (CNNs) have so far been the de-facto model for visual data.
Transfusion: Understanding transfer learning for medical imaging
M. Raghu, C. Zhang, J. Kleinberg, and S. Bengio · 1902
Earlier work this paper cites.
Rapid learning or feature reuse? towards understanding the effectiveness of maml
A. Raghu, M. Raghu, S. Bengio, and O. Vinyals · 1909
Earlier work this paper cites.
Similarity analysis of contextual word representation models
J. M. Wu, Y. Belinkov, H. Sajjad, N. Durrani, F. Dalvi, and J. Glass · 2005
Earlier work this paper cites.
Visual transformers: Token-based image representation and processing for computer vision
B. Wu, C. Xu, X. Dai, A. Wan, P. Zhang, Z. Yan, M. Tomizuka, J. Gonzalez, K. Keutzer, and P. Vajda · 2006
Earlier work this paper cites.
A kernel statistical test of independence
A. Gretton, K. Fukumizu, C. H. Teo, L. Song, B. Schölkopf, A. J. Smola, et al · 2007
Earlier work this paper cites.
Representational similarity analysis-connecting the branches of systems neuroscience
N. Kriegeskorte, M. Mur, and P. A. Bandettini · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Algorithms for learning kernels based on centered alignment
C. Cortes, M. Mohri, and A. Rostamizadeh · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Feature selection via dependence maximization
L. Song, A. Smola, A. Gretton, J. Bedo, and K. Borgwardt · 2012
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain and Y. Bengio · 2016
Earlier work this paper cites.
Understanding the effective receptive field in deep convolutional neural networks
W. Luo, Y. Li, R. Urtasun, and R. Zemel · 2017
Earlier work this paper cites.
M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
A. Conneau, G. Kruszewski, G. Lample, L. Barrault, and M. Baroni · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Insights on representational similarity in neural networks with canonical correlation
A. S. Morcos, M. Raghu, and S. Bengio · 2018
Cited alongside, same era.
Image transformer
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran · 2018
Cited alongside, same era.
Dissecting contextual word embeddings: Architecture and representation
M. E. Peters, M. Neumann, L. Zettlemoyer, and W.-t. Yih · 2018
Cited alongside, same era.
Attention augmented convolutional networks
I. Bello, B. Zoph, A. Vaswani, J. Shlens, and Q. V. Le · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Later among the works it cites.
What’s in a loss function for image classification?
S. Kornblith, H. Lee, T. Chen, and M. Norouzi · 2020
Later among the works it cites.
Convolutional neural networks as a model of the visual system: past, present, and future
G. W. Lindsay · 2020
Later among the works it cites.
What happens to bert embeddings during fine-tuning?
A. Merchant, E. Rahimtoroghi, E. Pavlick, and I. Tenney · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J.-B. Cordonnier, A. Loukas, and M. Jaggi · 2019
Cited alongside, same era.
Big transfer (bit): General visual representation learning
A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
S. Kornblith, M. Norouzi, H. Lee, and G. Hinton · 2019
Cited alongside, same era.
Investigating multilingual nmt representations at scale
S. R. Kudugunta, A. Bapna, I. Caswell, N. Arivazhagan, and O. Firat · 2019
Cited alongside, same era.
Universality and individuality in neural dynamics across large populations of recurrent networks
N. Maheswaranathan, A. H. Williams, M. D. Golub, S. Ganguli, and D. Sussillo · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens · 2019
Cited alongside, same era.
Comparison against task driven artificial neural networks reveals functional properties in mouse visual cortex
J. Shi, E. Shea-Brown, and M. Buice · 2019
Cited alongside, same era.
T. Nguyen, M. Raghu, and S. Kornblith · 2020
Later among the works it cites.
Understanding robustness of transformers for image classification
S. Bhojanapalli, A. Chakrabarti, D. Glasner, D. Li, T. Unterthiner, and A. Veit · 2021
Closest in time.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Closest in time.
An empirical study of training self-supervised vision transformers
X. Chen, S. Xie, and K. He · 2021
Closest in time.
Convit: Improving vision transformers with soft convolutional inductive biases
S. d’Ascoli, H. Touvron, M. Leavitt, A. Morcos, G. Biroli, and L. Sagun · 2021
Closest in time.
Supervised transfer learning at scale for medical imaging
B. Mustafa, A. Loh, J. Freyberg, P. MacWilliams, M. Wilson, S. M. McKinney, M. Sieniek, J. Winkens, Y. Liu, P. Bui, et al · 2021
Closest in time.
Intriguing properties of vision transformers, 2021
M. Naseer, K. Ranasinghe, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang · 2021
Closest in time.
Vision transformers are robust learners
S. Paul and P.-Y. Chen · 2021
Closest in time.
Are pre-trained convolutions better than pre-trained transformers?
Y. Tay, M. Dehghani, J. Gupta, D. Bahri, V. Aribandi, Z. Qin, and D. Metzler · 2021
Closest in time.
Mlp-mixer: An all-mlp architecture for vision
I. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Unterthiner, J. Yung, D. Keysers, J. Uszkoreit, M. Lucic, et al · 2021
Closest in time.
Resmlp: Feedforward networks for image classification with data-efficient training
H. Touvron, P. Bojanowski, M. Caron, M. Cord, A. El-Nouby, E. Grave, A. Joulin, G. Synnaeve, J. Verbeek, and H. Jégou · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, Z. Jiang, F. E. Tay, J. Feng, and S. Yan · 2021
Closest in time.