Fetching the paper…
Reading the bibliography…
The advancement of object detection (OD) in open-vocabulary and open-world scenarios is a critical challenge in computer vision.
Wider face and pedestrian challenge 2018: Methods and results
Loy, C.C., Lin, D., Ouyang, W., Xiong, Y., Yang, S., Huang, Q., Zhou, D., Xia, W., Li, Q., Luo, P., et al., 2019 · 1902
Earlier work this paper cites.
Zhou, X., Wang, D., Krähenbühl, P., 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V., 2019 · 1907
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L.H., Yatskar, M., Yin, D., Hsieh, C.J., Chang, K.W., X, Y, 2019 · 1908
Earlier work this paper cites.
Cross-dataset training for class increasing object detection
Yao, Y., Wang, Y., Guo, Y., Lin, J., Qin, H., Yan, J., 2020 · 2001
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee. pp. 248–255
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L., 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al., 2020 · 2010
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A., 2010 · 2010
Earlier work this paper cites.
Towards a category-extended object detector without relabeling or conflicts
Zhao, B., Chen, C., Xiao, W., Xiao, X., Ju, Q., Xia, S., 2020 · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 580–587
Girshick, R., Donahue, J., Darrell, T., Malik, J., 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context, in: European conference on computer vision, Springer. pp. 740–755
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L., 2014 · 2014
Earlier work this paper cites.
Fast r-cnn, in: Proceedings of the IEEE international conference on computer vision, pp. 1440–1448
Girshick, R., 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, in: Proceedings of the IEEE international conference on computer vision, pp. 2641–2649
Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., Lazebnik, S., 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., Sun, J., 2015 · 2015
Earlier work this paper cites.
Ssd: Single shot multibox detector, in: European conference on computer vision, Springer. pp. 21–37
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., Berg, A.C., 2016 · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 779–788
Redmon, J., Divvala, S., Girshick, R., Farhadi, A., 2016 · 2016
Earlier work this paper cites.
Wider face: A face detection benchmark, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5525–5533
Yang, S., Luo, P., Loy, C.C., Tang, X., 2016 · 2016
Earlier work this paper cites.
Mask r-cnn, in: Proceedings of the IEEE international conference on computer vision, pp. 2961–2969
He, K., Gkioxari, G., Dollár, P., Girshick, R., 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2018 · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning, in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2556–2565
Sharma, P., Ding, N., Goodman, S., Soricut, R., 2018 · 2018
Cited alongside, same era.
Lvis: A dataset for large vocabulary instance segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5356–5364
Scaling up visual and vision-language representation learning with noisy text supervision, in: International Conference on Machine Learning, PMLR. pp. 4904–4916
Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q., Sung, Y.H., Li, Z., Duerig, T., 2021 · 2021
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1780–1790
Kamath, A., Singh, M., LeCun, Y., Synnaeve, G., Misra, I., Carion, N., 2021 · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision, in: International Conference on Machine Learning, PMLR. pp. 5583–5594
Kim, W., Son, B., Kim, I., 2021 · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., Hoi, S.C.H., 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gupta, A., Dollar, P., Girshick, R., 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., Lee, S., 2019 · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection, in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 8430–8439
Shao, S., Li, Z., Zhang, T., Peng, C., Yu, G., Zhang, X., Li, J., Sun, J., 2019 · 2019
Cited alongside, same era.
End-to-end object detection with transformers, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, Springer. pp. 213–229
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S., 2020 · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning, in: European conference on computer vision, Springer. pp. 104–120
Chen, Y.C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., Liu, J., 2020 · 2020
Cited alongside, same era.
Large-scale adversarial training for vision-and-language representation learning
Gan, Z., Chen, Y.C., Li, L., Zhu, C., Cheng, Y., Liu, J., 2020 · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks, in: European Conference on Computer Vision, Springer. pp. 121–137
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., et al., 2020 · 2020
Cited alongside, same era.
Phrasecut: Language-based image segmentation in the wild, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10216–10225
Wu, C., Lin, Z., Cohen, S., Bui, T., Maji, S., 2020 · 2020
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10012–10022
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B., 2021 · 2021
Later among the works it cites.
Visualsparta: An embarrassingly simple approach to large-scale text-to-image search with weighted bag-of-words, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 5020–5029
Lu, X., Zhao, T., Lee, K., 2021 · 2021
Later among the works it cites.
Imagenet-21k pretraining for the masses
Ridnik, T., Ben-Baruch, E., Noy, A., Zelnik-Manor, L., 2021 · 2021
Later among the works it cites.
Sparse r-cnn: End-to-end object detection with learnable proposals, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14454–14463
Sun, P., Zhang, R., Jiang, Y., Kong, T., Xu, C., Zhan, W., Tomizuka, M., Li, L., Yuan, Z., Wang, C., et al., 2021 · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5579–5588
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., Gao, J., 2021 · 2021
Later among the works it cites.
Roboflow 100: A rich, multi-domain object detection benchmark
Ciaglia, F., Saverio Zuppichini, F., Guerrie, P., McQuade, M., Solawetz, J., 2022 · 2022
Closest in time.
Learning to prompt for open-vocabulary object detection with vision-language model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14084–14093
Du, Y., Wei, F., Zhang, Z., Shi, M., Gao, Y., Li, G., 2022 · 2022
Closest in time.
A convnet for the 2020s, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11976–11986
Liu, Z., Mao, H., Wu, C.Y., Feichtenhofer, C., Darrell, T., Xie, S., 2022 · 2022
Closest in time.
Detection hub: Unifying object detection datasets via query adaptation on language embedding
Meng, L., Dai, X., Chen, Y., Zhang, P., Chen, D., Liu, M., Wang, J., Wu, Z., Yuan, L., Jiang, Y.G., 2022 · 2022
Closest in time.
Simple open-vocabulary object detection with vision transformers
Minderer, M., Gritsenko, A., Stone, A., Neumann, M., Weissenborn, D., Dosovitskiy, A., Mahendran, A., Arnab, A., Dehghani, M., Shen, Z., et al., 2022 · 2022
Closest in time.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection
Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L.M., Shum, H.Y., 2022 · 2022
Closest in time.
An explainable toolbox for evaluating pre-trained vision-language models, in: Proceedings of the The 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 30–37
Zhao, T., Zhang, T., Zhu, M., Shen, H., Lee, K., Lu, X., Yin, J., 2022 · 2022
Closest in time.
Regionclip: Region-based language-image pretraining, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16793–16803
Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L.H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al., 2022 · 2022
Closest in time.