Fetching the paper…
Reading the bibliography…
Open-vocabulary learning has emerged as a cutting-edge research area, particularly in light of the widespread adoption of vision-based foundational models.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,”
2014
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Wang, M. Azab, N. Kojima, R. Mihalcea, and J. Deng, “Structured matching for phrase localization,” in
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko, “Modeling relationships in referential expressions with compositional modular networks,” in
2017
Earlier work this paper cites.
B. A. Plummer, A. Mallya, C. M. Cervantes, J. Hockenmaier, and S. Lazebnik, “Phrase localization and visual relationship detection with comprehensive image-language cues,” in
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
F. Zhao, J. Li, J. Zhao, and J. Feng, “Weakly supervised phrase localization with multi-scale anchored transformer network,” in
2018
Earlier work this paper cites.
Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high quality object detection,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Liu, H. Zhang, F. Wu, and Z.-J. Zha, “Learning to assemble neural module tree networks for visual grounding,” in
2019
Earlier work this paper cites.
R. Hong, D. Liu, X. Mo, X. He, and H. Zhang, “Learning to compose and reason with language tree structures for visual grounding,”
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Datta, K. Sikka, A. Roy, K. Ahuja, D. Parikh, and A. Divakaran, “Align2ground: Weakly supervised phrase grounding guided by image-caption alignment,” in
2019
Earlier work this paper cites.
J. Wang and L. Specia, “Phrase localization without paired training examples,” in
2019
Earlier work this paper cites.
K. Chen, J. Pang, J. Wang, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Shi, W. Ouyang, C. C. Loy, and D. Lin, “Hybrid task cascade for instance segmentation,” in
2019
Earlier work this paper cites.
A. Gupta, P. Dollar, and R. Girshick, “Lvis: A dataset for large vocabulary instance segmentation,” in
2019
Earlier work this paper cites.
Y. Liu, B. Wan, X. Zhu, and X. He, “Learning cross-modal context graph for visual grounding,” in
2020
Earlier work this paper cites.
B. Cheng, M. D. Collins, Y. Zhu, T. Liu, T. S. Huang, H. Adam, and L.-C. Chen, “Panoptic-DeepLab: A simple, strong, and fast baseline for bottom-up panoptic segmentation,” in
2020
Earlier work this paper cites.
X. Li, X. Li, L. Zhang, C. Guangliang, J. Shi, Z. Lin, Y. Tong, and S. Tan, “Improving semantic segmentation via decoupled body and edge supervision,” in
2020
Earlier work this paper cites.
X. Li, A. You, Z. Zhu, H. Zhao, M. Yang, K. Yang, and Y. Tong, “Semantic flow for fast and accurate scene parsing,” in
2020
Earlier work this paper cites.
X. Wang, R. Zhang, T. Kong, L. Li, and C. Shen, “SOLOv2: Dynamic and fast instance segmentation,” in
2020
Earlier work this paper cites.
J. Deng, Z. Yang, T. Chen
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Earlier work this paper cites.
X. Li, X. Li, A. You, L. Zhang, G.-L. Cheng, K. Yang, Y. Tong, and Z. Lin, “Towards efficient scene understanding via squeeze reasoning,”
2021
Earlier work this paper cites.
H. Wang, Y. Zhu, H. Adam, A. Yuille, and L.-C. Chen, “Max-deeplab: End-to-end panoptic segmentation with mask transformers,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Zareian, K. D. Rosa, D. H. Hu, and S.-F. Chang, “Open-vocabulary object detection using captions,” in
2021
Cited alongside, same era.
J. Wu, C. Wu, J. Lu, L. Wang, and X. Cui, “Region reinforcement network with topic constraint for image-text matching,”
2021
Cited alongside, same era.
P. Keserwani and P. P. Roy, “Text region conditional generative adversarial network for text concealment in the wild,”
2021
Cited alongside, same era.
G. Ghiasi, Y. Cui, A. Srinivas, R. Qian, T.-Y. Lin, E. D. Cubuk, Q. V. Le, and B. Zoph, “Simple copy-paste is a strong data augmentation method for instance segmentation,” in
2021
Cited alongside, same era.
H. Tan, X. Liu, B. Yin, and X. Li, “Cross-modal semantic matching generative adversarial networks for text-to-image synthesis,”
2021
Cited alongside, same era.
H. Li, M. Sun, J. Xiao, E. G. Lim, and Y. Zhao, “Fully and weakly supervised referring expression segmentation with end-to-end learning,”
2023
Closest in time.
M. Li, C. Wang, W. Feng, S. Lyu, G. Cheng, X. Li, B. Liu, and Q. Zhao, “Iterative robust visual grounding with masked reference based centerpoint supervision,”
2023
Closest in time.
Z. Wang, C. Yang, B. Jiang, and J. Yuan, “A dual reinforcement learning framework for weakly supervised phrase grounding,”
2023
Closest in time.
X. Li, H. Yuan, W. Zhang, G. Cheng, J. Pang, and C. C. Loy, “Tube-link: A flexible cross tube baseline for universal video segmentation,” in
2023
Closest in time.
L. Wang, Y. Liu, P. Du, Z. Ding, Y. Liao, Q. Qi, B. Chen, and S. Liu, “Object-aware distillation pyramid for open-vocabulary object detection,” in
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in
2021
Cited alongside, same era.
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion, “Mdetr-modulated detection for end-to-end multi-modal understanding,” in
2021
Cited alongside, same era.
Z. Fu, A. Kumar, A. Agarwal, and et al., “Coupling vision and proprioception for navigation of legged robots,” in
2022
Cited alongside, same era.
K. Sun, C. Guo, H. Zhang, and et al., “HVLM: exploring human-like visual cognition and language-memory network for visual dialog,”
2022
Cited alongside, same era.
L. Yang, Y. Xu, C. Yuan
2022
Cited alongside, same era.
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang
2022
Cited alongside, same era.
J. Wu, X. Li, X. Li, H. Ding, Y. Tong, and D. Tao, “Towards robust referring image segmentation,”
2022
Cited alongside, same era.
2023
Closest in time.
J. Qin, J. Wu, P. Yan, M. Li, R. Yuxi, X. Xiao, Y. Wang, R. Wang, S. Wen, X. Pan
2023
Closest in time.
X. Zou, Z.-Y. Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, H. Behl, J. Wang, L. Yuan
2023
Closest in time.
J. Wu, X. Li, S. X. H. Yuan, H. Ding, Y. Yang, X. Li, J. Zhang, Y. Tong, X. Jiang, B. Ghanem
2023
Closest in time.
Q. Zhao, S. Lyu, L. Chen, B. Liu, T.-B. Xu, G. Cheng, and W. Feng, “Learn by oneself: Exploiting weight-sharing potential in knowledge distillation guided ensemble network,”
2023
Closest in time.
2023
Closest in time.
P. Kaul, W. Xie, and A. Zisserman, “Multi-modal classifiers for open-vocabulary object detection,”
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
X. Wu, F. Zhu, R. Zhao, and H. Li, “Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching,” in
2023
Closest in time.
H. Song and J. Bang, “Prompt-guided transformers for end-to-end open-vocabulary object detection,”
2023
Closest in time.
W. Kuo, Y. Cui, X. Gu, A. Piergiovanni, and A. Angelova, “Open-vocabulary object detection upon frozen vision and language models,” in
2023
Closest in time.
L. Yao, J. Han, X. Liang, D. Xu, W. Zhang, Z. Li, and H. Xu, “Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment,” in
2023
Closest in time.
K. Han, Y. Liu, J. H. Liew, H. Ding, Y. Wei, J. Liu, Y. Wang, Y. Tang, Y. Yang, J. Feng
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
F. Liang, B. Wu, X. Dai, K. Li, Y. Zhao, H. Zhang, P. Zhang, P. Vajda, and D. Marculescu, “Open-vocabulary semantic segmentation with mask-adapted clip,” in
2023
Closest in time.
J. Wu, X. Li, H. Ding, X. Li, G. Cheng, Y. Tong, and C. C. Loy, “Betrayed by captions: Joint caption grounding and generation for open vocabulary instance segmentation,”
2023
Closest in time.
2023
Closest in time.
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, and S. De Mello, “Open-vocabulary panoptic segmentation with text-to-image diffusion models,” in
2023
Closest in time.
2023
Closest in time.
W. Wu, Y. Zhao, M. Z. Shou, H. Zhou, and C. Shen, “Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models,”
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.