Fetching the paper…
Reading the bibliography…
Transformers, which are popular for language modeling, have been explored for solving vision tasks recently, e.g., the Vision Transformer (ViT) for image classification.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis · 2016
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
T. DeVries and G. W. Taylor · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Earlier work this paper cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Residual attention network for image classification
F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, and X. Tang · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Gather-excite: Exploiting feature context in convolutional neural networks
J. Hu, L. Shen, S. Albanie, G. Sun, and A. Vedaldi · 2018
Earlier work this paper cites.
Squeeze-and-excitation networks
J. Hu, L. Shen, and G. Sun · 2018
Cited alongside, same era.
N. Parmar, A. Vaswani, J. Uszkoreit, Ł. Kaiser, N. Shazeer, A. Ku, and D. Tran · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training, 2018
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen · 2018
Cited alongside, same era.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Cited alongside, same era.
Cbam: Convolutional block attention module
S. Woo, J. Park, J.-Y. Lee, and I. So Kweon · 2018
Cited alongside, same era.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Later among the works it cites.
Pre-trained image processing transformer
H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao · 2020
Later among the works it cites.
Generative pretraining from pixels
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever · 2020
Later among the works it cites.
Rethinking attention with performers
K. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Davis, A. Mohiuddin, L. Kaiser, et al · 2020
Later among the works it cites.
Up-detr: Unsupervised pre-training for object detection with transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Psanet: Point-wise spatial attention network for scene parsing
H. Zhao, Y. Zhang, S. Liu, J. Shi, C. Change Loy, D. Lin, and J. Jia · 2018
Cited alongside, same era.
End-to-end dense video captioning with masked transformer
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong · 2018
Cited alongside, same era.
Attention augmented convolutional networks
I. Bello, B. Zoph, A. Vaswani, J. Shlens, and Q. V. Le · 2019
Cited alongside, same era.
Dual attention network for scene segmentation
J. Fu, J. Liu, H. Tian, Y. Li, Y. Bao, Z. Fang, and H. Lu · 2019
Cited alongside, same era.
Local relation networks for image recognition
H. Hu, Z. Zhang, Z. Xie, and S. Lin · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Z. Dai, B. Cai, Y. Lin, and J. Chen · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Later among the works it cites.
Ghostnet: More features from cheap operations
K. Han, Y. Wang, Q. Tian, J. Guo, C. Xu, and C. Xu · 2020
Later among the works it cites.
Rethinking transformer-based set prediction for object detection
Z. Sun, S. Cao, Y. Yang, and K. Kitani · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2020
Later among the works it cites.
End-to-end video instance segmentation with transformers
Y. Wang, Z. Xu, X. Wang, C. Shen, B. Cheng, H. Shen, and H. Xia · 2020
Later among the works it cites.
Visual transformers: Token-based image representation and processing for computer vision
B. Wu, C. Xu, X. Dai, A. Wan, P. Zhang, M. Tomizuka, K. Keutzer, and P. Vajda · 2020
Later among the works it cites.
Learning texture transformer network for image super-resolution
F. Yang, H. Yang, J. Fu, H. Lu, and B. Guo · 2020
Later among the works it cites.
A simple baseline for pose tracking in videos of crowed scenes
L. Yuan, S. Chang, Z. Huang, Y. Zhou, Y. Chen, X. Nie, F. E. Tay, J. Feng, and S. Yan · 2020
Later among the works it cites.
Revisiting knowledge distillation via label smoothing regularization
L. Yuan, F. E. Tay, G. Li, T. Wang, and J. Feng · 2020
Later among the works it cites.
Toward accurate person-level action recognition in videos of crowed scenes
L. Yuan, Y. Zhou, S. Chang, Z. Huang, Y. Chen, X. Nie, T. Wang, J. Feng, and S. Yan · 2020
Later among the works it cites.
Learning joint spatial-temporal transformations for video inpainting
Y. Zeng, J. Fu, and H. Chao · 2020
Later among the works it cites.
Exploring self-attention for image recognition
H. Zhao, J. Jia, and V. Koltun · 2020
Later among the works it cites.
H. Zhao, L. Jiang, J. Jia, P. Torr, and V. Koltun · 2020
Later among the works it cites.
End-to-end object detection with adaptive clustering transformer
M. Zheng, P. Gao, X. Wang, H. Li, and H. Dong · 2020
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2020
Later among the works it cites.