Fetching the paper…
Reading the bibliography…
Motivated by the remarkable achievements of DETR-based approaches on COCO object detection and segmentation benchmarks, recent endeavors have been directed towards elevating their performance through self-supervised pre-training of Transformers while preserving a frozen backbone.
Kuhn HW (1955) The hungarian method for the assignment problem. Naval research logistics quarterly 2(1-2):83–97
1955
Earlier work this paper cites.
Uijlings JR, Van De Sande KE, Gevers T, Smeulders AW (2013) Selective search for object recognition. IJCV 104(2):154–171
2013
Earlier work this paper cites.
Shao S, Li Z, Zhang T, Peng C, Yu G, Zhang X, Li J, Sun J (2019) Objects365: A large-scale, high-quality dataset for object detection. In: ICCV, pp 8430–8439
2019
Earlier work this paper cites.
Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A, Zagoruyko S (2020) End-to-end object detection with transformers. In: ECCV, Springer, pp 213–229
2020
Earlier work this paper cites.
Caron M, Misra I, Mairal J, Goyal P, Bojanowski P, Joulin A (2020) Unsupervised learning of visual features by contrasting cluster assignments. NeurIPS 33:9912–9924
2020
Earlier work this paper cites.
Chen T, Kornblith S, Norouzi M, Hinton G (2020) A simple framework for contrastive learning of visual representations. In: ICML, PMLR, pp 1597–1607
2020
Earlier work this paper cites.
Grill JB, Strub F, Altché F, Tallec C, Richemond P, Buchatskaya E, Doersch C, Avila Pires B, Guo Z, Gheshlaghi Azar M, et al. (2020) Bootstrap your own latent-a new approach to self-supervised learning. NeurIPS 33:21271–21284
2020
Earlier work this paper cites.
He K, Fan H, Wu Y, Xie S, Girshick R (2020) Momentum contrast for unsupervised visual representation learning. In: CVPR, pp 9729–9738
2020
Earlier work this paper cites.
Xie Q, Luong MT, Hovy E, Le QV (2020) Self-training with noisy student improves imagenet classification. In: CVPR, pp 10687–10698
2020
Earlier work this paper cites.
Zhu X, Su W, Lu L, Li B, Wang X, Dai J (2020) Deformable detr: Deformable transformers for end-to-end object detection. In: ICLR
2020
Earlier work this paper cites.
Zoph B, Ghiasi G, Lin TY, Cui Y, Liu H, Cubuk ED, Le Q (2020) Rethinking pre-training and self-training. NeurIPS 33:3833–3845
2020
Earlier work this paper cites.
Dai Z, Cai B, Lin Y, Chen J (2021) Up-detr: Unsupervised pre-training for object detection with transformers. In: CVPR, pp 1601–1610
2021
Cited alongside, same era.
Li Y, Wu CY, Fan H, Mangalam K, Xiong B, Malik J, Feichtenhofer C (2021) Improved multiscale vision transformers for classification and detection. arXiv preprint arXiv:211201526
2021
Cited alongside, same era.
Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B (2021) Swin transformer: Hierarchical vision transformer using shifted windows. In: ICCV, pp 10012–10022
2021
Cited alongside, same era.
Ranftl R, Bochkovskiy A, Koltun V (2021) Vision transformers for dense prediction. In: ICCV, pp 12179–12188
2021
Cited alongside, same era.
Wei F, Gao Y, Wu Z, Hu H, Lin S (2021) Aligning pretraining for detection via object-level contrastive learning. NeurIPS 34:22682–22694
2021
Wang F, Wang H, Wei C, Yuille A, Shen W (2022) Cp 2: Copy-paste contrastive pretraining for semantic segmentation. In: ECCV, Springer, pp 499–515
2022
Later among the works it cites.
Yang L, Zhuo W, Qi L, Shi Y, Gao Y (2022) St++: Make self-training work better for semi-supervised semantic segmentation. In: CVPR, pp 4268–4277
2022
Later among the works it cites.
Zhang H, Li F, Liu S, Zhang L, Su H, Zhu J, Ni LM, Shum HY (2022) Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv preprint arXiv:220303605
2022
Later among the works it cites.
Zong Z, Song G, Liu Y (2022) Detrs with collaborative hybrid assignments training. arXiv preprint arXiv:221112860
2022
Later among the works it cites.
Huang G, Li W, Teng J, Wang K, Chen Z, Shao J, Loy CC, Sheng L (2023) Siamese detr. In: CVPR, pp 15722–15731
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Zhu Y, Zhang Z, Wu C, Zhang Z, He T, Zhang H, Manmatha R, Li M, Smola AJ (2021) Improving semantic segmentation via efficient self-training. PAMI
2021
Cited alongside, same era.
Bar A, Wang X, Kantorov V, Reed CJ, Herzig R, Chechik G, Rohrbach A, Darrell T, Globerson A (2022) Detreg: Unsupervised pretraining with region priors for object detection. In: CVPR, pp 14605–14615
2022
Cited alongside, same era.
Chen Q, Wang J, Han C, Zhang S, Li Z, Chen X, Chen J, Wang X, Han S, Zhang G, et al. (2022) Group detr v2: Strong object detector with encoder-decoder pretraining. arXiv preprint arXiv:221103594
2022
Cited alongside, same era.
Li Z, Zhu Y, Yang F, Li W, Zhao C, Chen Y, Chen Z, Xie J, Wu L, Zhao R, et al. (2022) Univip: A unified framework for self-supervised visual pre-training. In: CVPR, pp 14627–14636
2022
Cited alongside, same era.
Sahito A, Frank E, Pfahringer B (2022) Better self-training for image classification through self-supervision. In: AJCAI, Springer, pp 645–657
2022
Cited alongside, same era.
Vandeghen R, Louppe G, Van Droogenbroeck M (2022) Adaptive self-training for object detection. arXiv preprint arXiv:221205911
2022
Cited alongside, same era.
Liu S, Li F, Zhang H, Yang X, Qi X, Su H, Zhu J, Zhang L (2022a) Dab-detr: Dynamic anchor boxes are better queries for detr. arXiv preprint arXiv:220112329
Cited in the paper.
Closest in time.
Jia D, Yuan Y, He H, Wu X, Yu H, Lin W, Sun L, Zhang C, Hu H (2023) Detrs with hybrid matching. In: CVPR, pp 19702–19712
2023
Closest in time.
Li F, Zhang H, Xu H, Liu S, Zhang L, Ni LM, Shum HY (2023) Mask dino: Towards a unified transformer-based framework for object detection and segmentation. In: CVPR, pp 3041–3050
2023
Closest in time.
Liu H, Li C, Wu Q, Lee YJ (2023) Visual instruction tuning
2023
Closest in time.
Podell D, English Z, Lacey K, Blattmann A, Dockhorn T, Müller J, Penna J, Rombach R (2023) Sdxl: Improving latent diffusion models for high-resolution image synthesis
2023
Closest in time.
Zhang L, Agrawala M (2023) Adding conditional control to text-to-image diffusion models. arXiv preprint arXiv:230205543
2023
Closest in time.