Fetching the paper…
Reading the bibliography…
This paper presents an improved DETR detector that maintains a "plain" nature: using a single-scale feature map and global cross-attention calculations without specific locality constraints, in contrast to previous leading DETR-based detectors that reintroduce architectural inductive biases of multi-scale and locality into the decoder.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Fast r-cnn
R. Girshick · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Earlier work this paper cites.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Cascade r-cnn: Delving into high quality object detection
Z. Cai and N. Vasconcelos · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Relation networks for object detection
H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Self-attention with relative position representations
P. Shaw, J. Uszkoreit, and A. Vaswani · 2018
Earlier work this paper cites.
Hybrid task cascade for instance segmentation
K. Chen, J. Pang, J. Wang, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Shi, W. Ouyang, et al · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov · 2019
Earlier work this paper cites.
Local relation networks for image recognition
H. Hu, Z. Zhang, Z. Xie, and S. Lin · 2019
Earlier work this paper cites.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Improve transformer models with better relative position embeddings
Z. Huang, D. Liang, P. Xu, and B. Xiang · 2020
Cited alongside, same era.
Efficientdet: Scalable and efficient object detection
M. Tan, R. Pang, and Q. V. Le · 2020
Cited alongside, same era.
Simmim: A simple framework for masked image modeling
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu · 2021
Later among the works it cites.
Group detr v2: Strong object detector with encoder-decoder pretraining
Q. Chen, J. Wang, C. Han, S. Zhang, Z. Li, X. Chen, J. Chen, X. Wang, S. Han, G. Zhang, et al · 2022
Later among the works it cites.
Eva: Exploring the limits of masked visual representation learning at scale
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao · 2022
Later among the works it cites.
Unleashing vanilla vision transformer with masked image modeling for object detection
Y. Fang, S. Yang, S. Wang, Y. Ge, Y. Shan, and X. Wang · 2022
Later among the works it cites.
Adamixer: A fast-converging query-based object detector
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raft: Recurrent all-pairs field transforms for optical flow
Z. Teed and J. Deng · 2020
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2020
Cited alongside, same era.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, S. Piao, and F. Wei · 2021
Cited alongside, same era.
Masked-attention mask transformer for universal image segmentation
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar · 2021
Cited alongside, same era.
Do we really need explicit position encodings for vision transformers
X. Chu, B. Zhang, Z. Tian, X. Wei, and H. Xia · 2021
Cited alongside, same era.
Dynamic detr: End-to-end object detection with dynamic attention
X. Dai, Y. Chen, J. Yang, P. Zhang, L. Yuan, and L. Zhang · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
Z. Gao, L. Wang, B. Han, and S. Guo · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Later among the works it cites.
D. Jia, Y. Yuan, H. He, X. Wu, H. Yu, W. Lin, L. Sun, C. Zhang, and H. Hu · 2022
Later among the works it cites.
Dn-detr: Accelerate detr training by introducing query denoising
F. Li, H. Zhang, S. Liu, J. Guo, L. M. Ni, and L. Zhang · 2022
Later among the works it cites.
Exploring plain vision transformer backbones for object detection
Y. Li, H. Mao, R. Girshick, and K. He · 2022
Later among the works it cites.
Dab-detr: Dynamic anchor boxes are better queries for detr
S. Liu, F. Li, H. Zhang, X. Yang, X. Qi, H. Su, J. Zhu, and L. Zhang · 2022
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution
Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, et al · 2022
Later among the works it cites.
J. Ouyang-Zhang, J. H. Cho, X. Zhou, and P. Krähenbühl · 2022
Later among the works it cites.
Internimage: Exploring large-scale vision foundation models with deformable convolutions
W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li, et al · 2022
Later among the works it cites.
Revealing the dark secrets of masked image modeling
Z. Xie, Z. Geng, J. Hu, Z. Zhang, H. Hu, and Y. Cao · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu · 2022
Later among the works it cites.
Semantic-aligned matching for enhanced detr convergence and multi-scale feature fusion
G. Zhang, Z. Luo, Y. Yu, J. Huang, K. Cui, S. Lu, and E. P. Xing · 2022
Later among the works it cites.
Towards efficient use of multi-scale features in transformer-based object detectors
G. Zhang, Z. Luo, Y. Yu, Z. Tian, J. Zhang, and S. Lu · 2022
Later among the works it cites.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, and H.-Y. Shum · 2022
Later among the works it cites.