Fetching the paper…
Reading the bibliography…
In this paper, we for the first time explore helpful multi-modal contextual knowledge to understand novel categories for open-vocabulary object detection (OVD).
I. Csiszár, “I-divergence geometry of probability distributions and minimization problems,” The Annals of Probability , pp. 146–158, 1975
1975
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” Transactions of the Association for Computational Linguistics , vol. 2, pp. 67–78, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Proceedings of the European Conference on Computer Vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” Advances in Neural Information Processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proceedings of the IEEE International Conference on Computer Vision , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Korattikara Balan, V. Rathod, K. P. Murphy, and M. Welling, “Bayesian dark knowledge,” Advances in Neural Information Processing Systems , vol. 28, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Schuster, R. Krishna, A. Chang, L. Fei-Fei, and C. D. Manning, “Generating semantically precise scene graphs from textual descriptions for improved image retrieval,” in Proceedings of the fourth workshop on vision and language , 2015, pp. 70–80
2015
Earlier work this paper cites.
P. Luo, Z. Zhu, Z. Liu, X. Wang, and X. Tang, “Face model compression by distilling knowledge from neurons,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 30, no. 1, 2016
2016
Earlier work this paper cites.
S. Gupta, J. Hoffman, and J. Malik, “Cross modal distillation for supervision transfer,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 2827–2836
2016
Earlier work this paper cites.
H. Bilen and A. Vedaldi, “Weakly supervised deep detection networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 2846–2854
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International Journal of Computer Vision , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6904–6913
2017
Earlier work this paper cites.
A. Bansal, K. Sikka, G. Sharma, R. Chellappa, and A. Divakaran, “Zero-shot object detection,” in Proceedings of the European Conference on Computer Vision , 2018, pp. 384–400
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of ACL , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6700–6709
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi, “From recognition to cognition: Visual commonsense reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6720–6731
2019
Earlier work this paper cites.
K. Li, Y. Zhang, K. Li, Y. Li, and Y. Fu, “Visual semantic reasoning for image-text matching,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 4654–4662
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
A. Gupta, P. Dollar, and R. Girshick, “Lvis: A dataset for large vocabulary instance segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 5356–5364
2019
Cited alongside, same era.
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun, “Objects365: A large-scale, high-quality dataset for object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 8430–8439
2019
Cited alongside, same era.
H. Wu, J. Mao, Y. Zhang, Y. Jiang, L. Li, W. Sun, and W.-Y. Ma, “Scenegraphparser,” https://github.com/vacancy/SceneGraphParser
2019
Cited alongside, same era.
Y. Xu, H. Wei, M. Lin, Y. Deng, K. Sheng, M. Zhang, F. Tang, W. Dong, F. Huang, and C. Xu, “Transformers in computational visual media: A survey,” Computational Visual Media , vol. 8, pp. 33–62, 2022
2022
Later among the works it cites.
J. Yang, J. Duan, S. Tran, Y. Xu, S. Chanda, L. Chen, B. Zeng, T. Chilimbi, and J. Huang, “Vision-language pre-training with triple contrastive learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 671–15 680
2022
Later among the works it cites.
2022
Later among the works it cites.
G. Luo, Y. Zhou, X. Sun, Y. Wang, L. Cao, Y. Wu, F. Huang, and R. Ji, “Towards lightweight transformer via group-wise transformation for vision-and-language tasks,” IEEE Transactions on Image Processing , vol. 31, pp. 3386–3398, 2022
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Li, N. Duan, Y. Fang, M. Gong, and D. Jiang, “Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2020, pp. 11 336–11 344
2020
Cited alongside, same era.
Y. Zhou, M. Wang, D. Liu, Z. Hu, and H. Zhang, “More grounded image captioning by distilling image-text matching model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 4777–4786
2020
Cited alongside, same era.
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov et al. , “The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale,” International Journal of Computer Vision , vol. 128, no. 7, pp. 1956–1981, 2020
2020
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 4904–4916
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Zareian, K. D. Rosa, D. H. Hu, and S.-F. Chang, “Open-vocabulary object detection using captions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 393–14 402
2021
Cited alongside, same era.
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” Advances in Neural Information Processing Systems , vol. 34, pp. 9694–9705, 2021
2021
Cited alongside, same era.
Later among the works it cites.
P. Zeng, H. Zhang, L. Gao, J. Song, and H. T. Shen, “Video question answering with prior knowledge and object-sensitive learning,” IEEE Transactions on Image Processing , vol. 31, pp. 5936–5948, 2022
2022
Later among the works it cites.
W. Zhao, Y. Rao, Y. Tang, J. Zhou, and J. Lu, “Videoabc: A real-world video dataset for abductive visual reasoning,” IEEE Transactions on Image Processing , vol. 31, pp. 6048–6061, 2022
2022
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 000–16 009
2022
Later among the works it cites.
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu, “Simmim: A simple framework for masked image modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 9653–9663
2022
Later among the works it cites.
Y. Gao, J.-X. Zhuang, S. Lin, H. Cheng, X. Sun, K. Li, and C. Shen, “Disco: Remedying self-supervised learning on lightweight models with distilled contrastive learning,” in European Conference on Computer Vision . Springer, 2022, pp. 237–253
2022
Later among the works it cites.
J. Song, Y. Chen, J. Ye, and M. Song, “Spot-adaptive knowledge distillation,” IEEE Transactions on Image Processing , vol. 31, pp. 3359–3370, 2022
2022
Later among the works it cites.
K. Li, J. Wan, and S. Yu, “Ckdf: Cascaded knowledge distillation framework for robust incremental learning,” IEEE Transactions on Image Processing , vol. 31, pp. 3825–3837, 2022
2022
Later among the works it cites.
Z. Huang, S. Yang, M. Zhou, Z. Li, Z. Gong, and Y. Chen, “Feature map distillation of thin nets for low-resolution object recognition,” IEEE Transactions on Image Processing , vol. 31, pp. 1364–1379, 2022
2022
Later among the works it cites.
S. Ge, B. Liu, P. Wang, Y. Li, and D. Zeng, “Learning privacy-preserving student networks via discriminative-generative distillation,” IEEE Transactions on Image Processing , vol. 32, pp. 116–127, 2022
2022
Later among the works it cites.
Z. Tu, X. Liu, and X. Xiao, “A general dynamic knowledge distillation method for visual analytics,” IEEE Transactions on Image Processing , vol. 31, pp. 6517–6531, 2022
2022
Later among the works it cites.
X. Wu, D. Hong, and J. Chanussot, “Uiu-net: U-net in u-net for infrared small object detection,” IEEE Transactions on Image Processing , vol. 32, pp. 364–376, 2022
2022
Later among the works it cites.
Z. Ma, G. Luo, J. Gao, L. Li, Y. Chen, S. Wang, C. Zhang, and W. Hu, “Open-vocabulary one-stage detection with hierarchical visual-language knowledge distillation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 14 074–14 083
2022
Later among the works it cites.
M. Gao, C. Xing, J. C. Niebles, J. Li, R. Xu, W. Liu, and C. Xiong, “Open vocabulary object detection with pseudo bounding-box labels,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part X . Springer, 2022, pp. 266–282
2022
Later among the works it cites.
C. Feng, Y. Zhong, Z. Jie, X. Chu, H. Ren, X. Wei, W. Xie, and L. Ma, “Promptdet: Towards open-vocabulary detection using uncurated images,” in Proceedings of the European Conference on Computer Vision , 2022
2022
Later among the works it cites.
S. Zhao, Z. Zhang, S. Schulter, L. Zhao, B. Vijay Kumar, A. Stathopoulos, M. Chandraker, and D. N. Metaxas, “Exploiting unlabeled data with vision and language models for object detection,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX . Springer, 2022, pp. 159–175
2022
Later among the works it cites.
Y. Du, F. Wei, Z. Zhang, M. Shi, Y. Gao, and G. Li, “Learning to prompt for open-vocabulary object detection with vision-language model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022
2022
Later among the works it cites.
2023
Closest in time.
J. Yang, X. Li, M. Zheng, Z. Wang, Y. Zhu, X. Guo, Y. Yuan, Z. Chai, and S. Jiang, “Membridge: Video-language pre-training with memory-augmented inter-modality bridge,” IEEE Transactions on Image Processing , 2023
2023
Closest in time.
X. Zhang, F. Zhang, and C. Xu, “Reducing vision-answer biases for multiple-choice vqa,” IEEE Transactions on Image Processing , 2023
2023
Closest in time.
Z. Li, Y. Guo, K. Wang, Y. Wei, L. Nie, and M. Kankanhalli, “Joint answering and explanation for visual commonsense reasoning,” IEEE Transactions on Image Processing , 2023
2023
Closest in time.