Fetching the paper…
Reading the bibliography…
Object detection (OD) in computer vision has made significant progress in recent years, transitioning from closed-set labels to open-vocabulary detection (OVD) based on large-scale vision-language pre-training (VLP).
Beyond accuracy: Behavioral testing of NLP models with CheckList
Ribeiro, M. T.; Wu, T.; Guestrin, C.; and Singh, S. 2020 · 2005
Earlier work this paper cites.
Pascal VOC 2008 challenge
Hoiem, D.; Divvala, S. K.; and Hays, J. H. 2009 · 2008
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Fast r-cnn
Girshick, R. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Redmon, J.; Divvala, S.; Girshick, R.; and Farhadi, A. 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Learning to detect human-object interactions
Chao, Y.-W.; Liu, Y.; Liu, X.; Zeng, H.; and Deng, J. 2018 · 2018
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
Gupta, A.; Dollar, P.; and Girshick, R. 2019 · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
Shao, S.; Li, Z.; Zhang, T.; Peng, C.; Yu, G.; Zhang, X.; Li, J.; and Sun, J. 2019 · 2019
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X.; Yin, X.; Li, C.; Zhang, P.; Hu, X.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; et al. 2020 · 2020
Cited alongside, same era.
Ppdm: Parallel point detection and matching for real-time human-object interaction detection
Liao, Y.; Liu, S.; Wang, F.; Chen, Y.; Qian, C.; and Feng, J. 2020 · 2020
Cited alongside, same era.
Phrasecut: Language-based image segmentation in the wild
Wu, C.; Lin, Z.; Cohen, S.; Bui, T.; and Maji, S. 2020 · 2020
Cited alongside, same era.
Mdetr-modulated detection for end-to-end multi-modal understanding
X-detr: A versatile architecture for instance-wise vision-language tasks
Cai, Z.; Kwon, G.; Ravichandran, A.; Bas, E.; Tu, Z.; Bhotika, R.; and Soatto, S. 2022 · 2022
Later among the works it cites.
Coarse-to-fine vision-language pre-training with fusion in the backbone
Dou, Z.-Y.; Kamath, A.; Gan, Z.; Zhang, P.; Wang, J.; Li, L.; Liu, Z.; Liu, C.; LeCun, Y.; Peng, N.; et al. 2022 · 2022
Later among the works it cites.
Du, X.; Legastelois, B.; Ganesh, B.; Rajan, A.; Chockler, H.; Belle, V.; Anderson, S.; and Ramamoorthy, S. 2022 · 2022
Later among the works it cites.
Detecting twenty-thousand classes using image-level supervision
Zhou, X.; Girdhar, R.; Joulin, A.; Krähenbühl, P.; and Misra, I. 2022 · 2022
Later among the works it cites.
Otter: A multi-modal model with in-context instruction tuning
Li, B.; Zhang, Y.; Chen, L.; Wang, J.; Yang, J.; and Liu, Z. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kamath, A.; Singh, M.; LeCun, Y.; Synnaeve, G.; Misra, I.; and Carion, N. 2021 · 2021
Cited alongside, same era.
Learning To Predict Visual Attributes in the Wild
Pham, K.; Kafle, K.; Lin, Z.; Ding, Z.; Cohen, S.; Tran, Q.; and Shrivastava, A. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models
Li, C.; Liu, H.; Li, L. H.; Zhang, P.; Aneja, J.; Yang, J.; Jin, P.; Hu, H.; Liu, Z.; Lee, Y. J.; and Gao, J. 2022a
Cited in the paper.
Grounded language-image pre-training
Li, L. H.; Zhang, P.; Zhang, H.; Yang, J.; Li, C.; Zhong, Y.; Wang, L.; Yuan, L.; Zhang, L.; Hwang, J.-N.; et al. 2022b
Cited in the paper.
Omdet: Language-aware object detection with large-scale vision-language multi-dataset pre-training
Zhao, T.; Liu, P.; Lu, X.; and Lee, K. 2022a
Cited in the paper.
Closest in time.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Li, C.; Yang, J.; Su, H.; Zhu, J.; et al. 2023 · 2023
Closest in time.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Shah, D.; Osiński, B.; Levine, S.; et al. 2023 · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; and Zhuang, Y. 2023 · 2023
Closest in time.
V3det: Vast vocabulary visual detection dataset
Wang, J.; Zhang, P.; Chu, T.; Cao, Y.; Zhou, Y.; Wu, T.; Wang, B.; He, C.; and Lin, D. 2023 · 2023
Closest in time.