Fetching the paper…
Reading the bibliography…
Transformers with remarkable global representation capacities achieve competitive results for visual tasks, but fail to consider high-level local pattern information in input images.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Image segmentation fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
S. Ren, K. He, R. B. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
SSD: single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C. Fu, and A. C. Berg · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi · 2016
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
F. Chollet · 2017
Earlier work this paper cites.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
T. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, and A. N. G. et al · 2017
Earlier work this paper cites.
Cascade r-cnn: Delving into high quality object detection
Z. Cai and N. Vasconcelos · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Squeeze-and-excitation networks
J. Hu, L. Shen, and G. Sun · 2018
Earlier work this paper cites.
Cornernet: Detecting objects as paired keypoints
H. Law and J. Deng · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Cited alongside, same era.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
X. Zhang, X. Zhou, M. Lin, and J. Sun · 2018
Cited alongside, same era.
Centernet: Object detection with keypoint triplets
K. Duan, S. Bai, L. Xie, H. Qi, Q. Huang, and Q. Tian · 2019
Cited alongside, same era.
Dynamic fusion with intra- and inter-modality attention flow for visual question answering
P. Gao, Z. Jiang, H. You, P. Lu, S. C. H. Hoi, X. Wang, and H. Li · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Twins: Revisiting the design of spatial attention in vision transformers
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, and X. W. et al · 2021
Closest in time.
Conditional positional encodings for vision transformers
X. Chu, Z. Tian, B. Zhang, X. Wang, X. Wei, and H. Xia · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, and T. U. et al · 2021
Closest in time.
Clip-adapter: Better vision-language models with feature adapters
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao · 2021
Closest in time.
Container: Context aggregation network
P. Gao, J. Lu, H. Li, R. Mottaghi, and A. Kembhavi · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
Libra r-cnn: Towards balanced learning for object detection
J. Pang, K. Chen, J. Shi, H. Feng, W. Ouyang, and D. Lin · 2019
Cited alongside, same era.
Multi-modality latent interaction network for visual question answering
G. Peng, H. You, Z. Zhang, X. Wang, and H. Li · 2019
Cited alongside, same era.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Cited alongside, same era.
FreeAnchor: Learning to match anchors for visual object detection
X. Zhang, F. Wan, C. Liu, R. Ji, and Q. Ye · 2019
Cited alongside, same era.
Asymmetric non-local neural networks for semantic segmentation
Z. Zhu, M. Xu, S. Bai, T. Huang, and X. Bai · 2019
Cited alongside, same era.
Fast convergence of detr with spatially modulated co-attention
P. Gao, M. Zheng, X. Wang, J. Dai, and H. Li · 2021
Closest in time.
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang · 2021
Closest in time.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, and Z. Z. et al · 2021
Closest in time.
Conformer: Local features coupling global representations for visual recognition
Z. Peng, W. Huang, S. Gu, L. Xie, Y. Wang, J. Jiao, and Q. Ye · 2021
Closest in time.
MaX-DeepLab: End-to-end panoptic segmentation with mask transformers
H. Wang, Y. Zhu, H. Adam, A. Yuille, and L.-C. Chen · 2021
Closest in time.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, and D. L. et al · 2021
Closest in time.
Cvt: Introducing convolutions to vision transformers
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, and L. Y. et al · 2021
Closest in time.
Contnet: Why not use convolution and transformer at the same time?
H. Yan, Z. Li, W. Li, C. Wang, M. Wu, and C. Zhang · 2021
Closest in time.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, and Z. J. et al · 2021
Closest in time.
Rest: An efficient transformer for visual recognition
Q. Zhang and Y. Yang · 2021
Closest in time.
Proto: Program-guided transformer for program-guided tasks
Z. Zhao, K. Samel, B. Chen, and L. Song · 2021
Closest in time.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, and Y. W. et al · 2021
Closest in time.
Deepvit: Towards deeper vision transformer
D. Zhou, B. Kang, X. Jin, L. Yang, X. Lian, and Z. J. et al · 2021
Closest in time.