Fetching the paper…
Reading the bibliography…
Although convolutional networks (ConvNets) have enjoyed great success in computer vision (CV), it suffers from capturing global information crucial to dense prediction tasks such as object detection and segmentation.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Optimization for machine learning
S. Sra, S. Nowozin, and S. J. Wright · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Human pose estimation via deep neural networks’
A. Toshev and C. Szegedy · 2014
Earlier work this paper cites.
Region-based convolutional networks for accurate object detection and segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis · 2016
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
F. Chollet · 2017
Earlier work this paper cites.
Deformable convolutional networks
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollar, and R. Girshick · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Cited alongside, same era.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Cited alongside, same era.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi · 2017
Cited alongside, same era.
Cbam: Convolutional block attention module
S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon · 2018
Later among the works it cites.
Attention augmented convolutional networks
I. Bello, B. Zoph, A. Vaswani, J. Shlens, and Q. V. Le · 2019
Later among the works it cites.
Gcnet: Non-local networks meet squeeze-excitation networks and beyond
Y. Cao, J. Xu, S. Lin, F. Wei, and H. Hu · 2019
Later among the works it cites.
Res2net: A new multi-scale backbone architecture
S. Gao, M.-M. Cheng, K. Zhao, X.-Y. Zhang, M.-H. Yang, and P. H. Torr · 2019
Later among the works it cites.
Bag of tricks for image classification with convolutional neural networks
T. He, Z. Zhang, H. Zhang, Z. Zhang, J. Xie, and M. Li · 2019
Later among the works it cites.
Local relation networks for image recognition
H. Hu, Z. Zhang, Z. Xie, and S. Lin · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Residual attention network for image classification
F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, and X. Tang · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Cited alongside, same era.
Pyramid scene parsing network
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia · 2017
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
E. D. Cubuk, B. Zoph, D. Mane, V. Vasudevan, and Q. V. Le · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Do better imagenet models transfer better?
S. Kornblith, J. Shlens, and Q. V. Le · 2019
Later among the works it cites.
Stand-alone self-attention in vision models
P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens · 2019
Later among the works it cites.
Fcos: Fully convolutional one-stage object detection
Z. Tian, C. Shen, H. Chen, and T. He · 2019
Later among the works it cites.
Deformable convnets v2: More deformable, better results
X. Zhu, H. Hu, S. Lin, and J. Dai · 2019
Later among the works it cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Later among the works it cites.
Convtransformer: A convolutional transformer network for video frame synthesis
Z. Liu, S. Luo, W. Li, J. Lu, Y. Wu, C. Li, and L. Yang · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou · 2020
Later among the works it cites.
Exploring self-attention for image recognition
H. Zhao, J. Jia, and V. Koltun · 2020
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr, et al · 2020
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2020
Later among the works it cites.
Transgan: Two transformers can make one strong gan
Y. Jiang, S. Chang, and Z. Wang · 2021
Closest in time.
Contnet: Why not use convolution and transformer at the same time?
H. Yan, Z. Li, W. Li, C. Wang, M. Wu, and C. Zhang · 2021
Closest in time.