Fetching the paper…
Reading the bibliography…
This report introduces an enhanced method for the Foundational Few-Shot Object Detection (FSOD) task, leveraging the vision-language model (VLM) for object detection.
Complex object classification: A multi-modal multi-instance multi-label deep network with optimal transport
Y. Yang, Y. Wu, D. Zhan, Z. Liu, and Y. Jiang · 2018
Earlier work this paper cites.
Comprehensive semi-supervised multi-modal learning
Y. Yang, K. Wang, D. Zhan, H. Xiong, and Y. Jiang · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom · 2020
Earlier work this paper cites.
Semi-supervised multi-modal multi-instance multi-label deep network with optimal transport
Y. Yang, Z. Fu, D. Zhan, Z. Liu, and Y. Jiang · 2021
Earlier work this paper cites.
S2osc: A holistic semi-supervised approach for open set classification
Y. Yang, H. Wei, Z.-Q. Sun, G.-Y. Li, Y. Zhou, H. Xiong, and J. Yang · 2021
Earlier work this paper cites.
Grounded language-image pre-training
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J. Hwang, K. Chang, and J. Gao · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Grounding DINO: marrying DINO with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, and L. Zhang · 2023
Cited alongside, same era.
Revisiting few-shot object detection with vision-language models
A. Madan, N. Peri, S. Kong, and D. Ramanan · 2023
Cited alongside, same era.
Robust semi-supervised learning for self-learning open-world classes
W. Xi, X. Song, W. Guo, and Y. Yang · 2023
Later among the works it cites.
Towards global video scene segmentation with context-aware transformer
Y. Yang, Y. Huang, W. Guo, B. Xu, and D. Xia · 2023
Later among the works it cites.
Learning to rebalance multi-modal optimization by adaptively masking subnetworks
Y. Yang, H. Pan, Q.-Y. Jiang, Y. Xu, and J. Tang · 2024
Closest in time.
Mm-llms: Recent advances in multimodal large language models
D. Zhang, Y. Yu, C. Li, J. Dong, D. Su, C. Chu, and D. Yu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…