Fetching the paper…
Reading the bibliography…
Convolutional architectures have proven extremely successful for vision tasks.
The need for biases in learning generalizations
Mitchell, T. M · 1980
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D · 1989
Earlier work this paper cites.
Statistics of natural images: Scaling in the woods
Ruderman, D. L. and Bialek, W · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Gers, F. A., Schmidhuber, J., and Cummins, F · 1999
Earlier work this paper cites.
Natural image statistics and neural representation
Simoncelli, E. P. and Olshausen, B. A · 2001
Earlier work this paper cites.
Visual Transformers: Token-based Image Representation and Processing for Computer Vision
Wu, B., Xu, C., Dai, X., Wan, A., Zhang, P., Tomizuka, M., Keutzer, K., and Vajda, P · 2006
Earlier work this paper cites.
Evaluation of Pooling Operations in Convolutional Architectures for Object Recognition
Scherer, D., Müller, A., and Behnke, S · 2010
Earlier work this paper cites.
LSTM neural networks for language modeling
Sundermeyer, M., Schlüter, R., and Ney, H · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Deep learning in neural networks: An overview
Schmidhuber, J · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Homotopy analysis for tensor pca
Anandkumar, A., Deng, Y., Ge, R., and Mobahi, H · 2016
Earlier work this paper cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
LSTM: A Search Space Odyssey
Greff, K., Srivastava, R. K., Koutník, J., Steunebrink, B. R., and Schmidhuber, J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
A2-nets: Double attention networks
Chen, Y., Kalantidis, Y., Li, J., Yan, S., and Feng, J · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Squeeze-and-Excitation Networks
Transferring inductive biases through knowledge distillation
Abnar, S., Dehghani, M., and Zuidema, W · 2020
Later among the works it cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Later among the works it cites.
Uniter: Universal image-text representation learning
Chen, Y.-C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Revisiting spatial invariance with low-rank local connectivity
Elsayed, G., Ramachandran, P., Shlens, J., and Kornblith, S · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, J., Shen, L., and Sun, G · 2018
Cited alongside, same era.
Non-local Neural Networks
Wang, X., Girshick, R., Gupta, A., and He, K · 2018
Cited alongside, same era.
Attention augmented convolutional networks
Bello, I., Zoph, B., Vaswani, A., Shlens, J., and Le, Q. V · 2019
Cited alongside, same era.
On the relationship between self-attention and convolutional layers
Cordonnier, J.-B., Loukas, A., and Jaggi, M · 2019
Cited alongside, same era.
Finding the needle in the haystack with convolutions: on the benefits of architectural bias
d’Ascoli, S., Sagun, L., Biroli, G., and Bruna, J · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J · 2019
Cited alongside, same era.
Adaptive attention span in transformers
Sukhbaatar, S., Grave, E., Bojanowski, P., and Joulin, A · 2019
Cited alongside, same era.
Later among the works it cites.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Later among the works it cites.
Towards learning convolutions from scratch
Neyshabur, B · 2020
Later among the works it cites.
Designing network design spaces
Radosavovic, I., Kosaraju, R. P., Girshick, R., He, K., and Dollár, P · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2020
Later among the works it cites.
Splitnet: Divide and co-training
Zhao, S., Zhou, L., Wang, W., Cai, D., Lam, T. L., and Xu, Y · 2020
Later among the works it cites.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong, Y., Cordonnier, J.-B., and Loukas, A · 2021
Closest in time.
Bottleneck Transformers for Visual Recognition
Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A · 2021
Closest in time.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Keysers, D., Uszkoreit, J., Lucic, M., et al · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Tay, F. E., Feng, J., and Yan, S · 2021
Closest in time.