Fetching the paper…
Reading the bibliography…
Transformers are successfully applied to computer vision due to their powerful modeling capacity with self-attention.
The need for biases in learning generalizations
Tom M Mitchell · 1980
Earlier work this paper cites.
Statistics of natural images: Scaling in the woods
Daniel L Ruderman and William Bialek · 1994
Earlier work this paper cites.
Evaluation and selection of biases in machine learning
Diana F Gordon and Marie Desjardins · 1995
Earlier work this paper cites.
Natural image statistics and neural representation
Eero P Simoncelli and Bruno A Olshausen · 2001
Earlier work this paper cites.
Automated flower classification over a large number of classes
M. E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Distance metric learning for large margin nearest neighbor classification
Kilian Q Weinberger and Lawrence K Saul · 2009
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and F. F. Li · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Inceptionism: Going deeper into neural networks
A. Mordvintsev, Christopher Olah, and Mike Tyka · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Mode regularized generative adversarial networks
Tong Che, Yanran Li, Athul Paul Jacob, Yoshua Bengio, and Wenjie Li · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
In defense of the triplet loss for person re-identification
Alexander Hermans, Lucas Beyer, and Bastian Leibe · 2017
Earlier work this paper cites.
Data-free knowledge distillation for deep neural networks
Raphael Gontijo Lopes, Stefano Fenu, and Thad Starner · 2017
Earlier work this paper cites.
Veegan: Reducing mode collapse in gans using implicit variational learning
Akash Srivastava, Lazar Valkov, Chris Russell, Michael U Gutmann, and Charles Sutton · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Earlier work this paper cites.
Random erasing data augmentation
Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Diverse image-to-image translation via disentangled representations
Hsin-Ying Lee, Hung-Yu Tseng, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2018
Cited alongside, same era.
Knowledge transfer with jacobian matching
Suraj Srinivas and François Fleuret · 2018
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Cited alongside, same era.
Dream distillation: A data-independent model compression framework
Designing network design spaces
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár · 2020
Later among the works it cites.
Improved embeddings with easy positive triplet mining
Hong Xuan, Abby Stylianou, and Robert Pless · 2020
Later among the works it cites.
Data-free knowledge amalgamation via group-stack dual-gan
Jingwen Ye, Yixin Ji, Xinchao Wang, Xin Gao, and Mingli Song · 2020
Later among the works it cites.
Dreaming to distill: Data-free knowledge transfer via deepinversion
Hongxu Yin, Pavlo Molchanov, Jose M Alvarez, Zhizhong Li, Arun Mallya, Derek Hoiem, Niraj K Jha, and Jan Kautz · 2020
Later among the works it cites.
Data-free knowledge distillation for object detection
Akshay Chawla, Hongxu Yin, Pavlo Molchanov, and Jose Alvarez · 2021
Later among the works it cites.
Pre-trained image processing transformer
Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kartikeya Bhardwaj, Naveen Suda, and Radu Marculescu · 2019
Cited alongside, same era.
Data-free learning of student networks
Hanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang, Chuanjian Liu, Boxin Shi, Chunjing Xu, Chao Xu, and Qi Tian · 2019
Cited alongside, same era.
Knowledge transfer via distillation of activation boundaries formed by hidden neurons
Byeongho Heo, Minsik Lee, Sangdoo Yun, and Jin Young Choi · 2019
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Cited alongside, same era.
Structured knowledge distillation for semantic segmentation
Yifan Liu, Ke Chen, Chris Liu, Zengchang Qin, Zhenbo Luo, and Jingdong Wang · 2019
Cited alongside, same era.
Zero-shot knowledge transfer via adversarial belief matching
Paul Micaelli and Amos Storkey · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Later among the works it cites.
Oh-former: Omni-relational high-order transformer for person re-identification
Xianing Chen, Jialang Xu, Jiale Xu, and Shenghua Gao · 2021
Later among the works it cites.
Coatnet: Marrying convolution and attention for all data sizes
Zihang Dai, Hanxiao Liu, Quoc V Le, and Mingxing Tan · 2021
Later among the works it cites.
Convit: Improving vision transformers with soft convolutional inductive biases
Stéphane d’Ascoli, Hugo Touvron, Matthew Leavitt, Ari Morcos, Giulio Biroli, and Levent Sagun · 2021
Later among the works it cites.
Levit: a vision transformer in convnet’s clothing for faster inference
Ben Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze · 2021
Later among the works it cites.
Rethinking spatial dimensions of vision transformers
Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh · 2021
Later among the works it cites.
Contextual transformer networks for visual recognition
Yehao Li, Ting Yao, Yingwei Pan, and Tao Mei · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Scalable vision transformers with hierarchical pooling
Zizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He, and Jianfei Cai · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Cvt: Introducing convolutions to vision transformers
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Later among the works it cites.
Early convolutions help transformers see better
Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollár, and Ross Girshick · 2021
Later among the works it cites.
Vitae: Vision transformer advanced by exploring intrinsic inductive bias
Yufei Xu, Qiming Zhang, Jing Zhang, and Dacheng Tao · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
L. Yuan, Y Chen, T. Wang, W. Yu, Y Shi, F. E. Tay, J. Feng, and S. Yan · 2021
Later among the works it cites.
Data-free knowledge distillation for image super-resolution
Yiman Zhang, Hanting Chen, Xinghao Chen, Yiping Deng, Chunjing Xu, and Yunhe Wang · 2021
Later among the works it cites.
Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond
Qiming Zhang, Yufei Xu, Jing Zhang, and Dacheng Tao · 2022
Closest in time.