Fetching the paper…
Reading the bibliography…
Understanding the semantics of individual regions or patches of unconstrained images, such as open-world object detection, remains a critical yet challenging task in computer vision.
Describing objects by their attributes
A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth · 2009
Earlier work this paper cites.
Evaluating knowledge transfer and zero-shot learning in a large-scale setting
M. Rohrbach, M. Stark, and B. Schiele · 2011
Earlier work this paper cites.
Learning to share visual appearance for multiclass object detection
R. Salakhutdinov, A. Torralba, and J. Tenenbaum · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov · 2013
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
R. Socher, M. Ganjoo, C. D. Manning, and A. Ng · 2013
Earlier work this paper cites.
Zero-shot recognition with unreliable attributes
D. Jayaraman and K. Grauman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Multi-cue zero-shot learning with strong supervision
Z. Akata, M. Malinowski, M. Fritz, and B. Schiele · 2016
Earlier work this paper cites.
Learning deep representations of fine-grained visual descriptions
S. Reed, Z. Akata, H. Lee, and B. Schiele · 2016
Earlier work this paper cites.
Predicting visual exemplars of unseen classes for zero-shot learning
S. Changpinyo, W.-L. Chao, and F. Sha · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Openimages: A public dataset for large-scale multi-label and multi-class image classification
I. Krasin, T. Duerig, N. Alldrin, V. Ferrari, S. Abu-El-Haija, A. Kuznetsova, H. Rom, J. Uijlings, S. Popov, A. Veit, et al · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Open vocabulary scene parsing
H. Zhao, X. Puig, B. Zhou, S. Fidler, and A. Torralba · 2017
Cited alongside, same era.
Zero-shot recognition via semantic embeddings and knowledge graphs
X. Wang, Y. Ye, and A. Gupta · 2018
Cited alongside, same era.
Lvis: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollar, and R. Girshick · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Cited alongside, same era.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Cited alongside, same era.
Grounded language-image pre-training
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, et al · 2022
Later among the works it cites.
Learning object-language alignments for open-vocabulary object detection
C. Lin, P. Sun, Y. Jiang, P. Luo, L. Qu, G. Haffari, Z. Yuan, and J. Cai · 2022
Later among the works it cites.
Open-vocabulary one-stage detection with hierarchical visual-language knowledge distillation
Z. Ma, G. Luo, J. Gao, L. Li, Y. Chen, S. Wang, C. Zhang, and W. Hu · 2022
Later among the works it cites.
Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection
L. Yao, J. Han, Y. Wen, X. Liang, D. Xu, W. Zhang, Z. Li, C. Xu, and H. Xu · 2022
Later among the works it cites.
Open-vocabulary detr with conditional matching
Y. Zang, W. Li, K. Zhou, C. Huang, and C. C. Loy · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Region graph embedding network for zero-shot learning
G.-S. Xie, L. Liu, F. Zhu, F. Zhao, Z. Zhang, Y. Yao, J. Qin, and L. Shao · 2020
Cited alongside, same era.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Mdetr-modulated detection for end-to-end multi-modal understanding
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Sparse r-cnn: End-to-end object detection with learnable proposals
P. Sun, R. Zhang, Y. Jiang, T. Kong, C. Xu, W. Zhan, M. Tomizuka, L. Li, Z. Yuan, C. Wang, et al · 2021
Cited alongside, same era.
Glipv2: Unifying localization and vision-language understanding
H. Zhang, P. Zhang, X. Hu, Y.-C. Chen, L. Li, X. Dai, L. Wang, L. Yuan, J.-N. Hwang, and J. Gao · 2022
Later among the works it cites.
Regionclip: Region-based language-image pretraining
Y. Zhong, J. Yang, P. Zhang, C. Li, N. Codella, L. H. Li, L. Zhou, X. Dai, L. Yuan, Y. Li, et al · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models
K. Zhou, J. Yang, C. C. Loy, and Z. Liu · 2022
Later among the works it cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Closest in time.
Codet: Co-occurrence guided region-word alignment for open-vocabulary object detection
C. Ma, Y. Jiang, X. Wen, Z. Yuan, and X. Qi · 2023
Closest in time.
V3det: Vast vocabulary visual detection dataset
J. Wang, P. Zhang, T. Chu, Y. Cao, Y. Zhou, T. Wu, B. Wang, C. He, and D. Lin · 2023
Closest in time.
Universal instance perception as object discovery and retrieval
B. Yan, Y. Jiang, J. Wu, D. Wang, P. Luo, Z. Yuan, and H. Lu · 2023
Closest in time.
Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment
L. Yao, J. Han, X. Liang, D. Xu, W. Zhang, Z. Li, and H. Xu · 2023
Closest in time.
A simple framework for open-vocabulary segmentation and detection
H. Zhang, F. Li, X. Zou, S. Liu, C. Li, J. Gao, J. Yang, and L. Zhang · 2023
Closest in time.
Generalized decoding for pixel, image, and language
X. Zou, Z.-Y. Dou, J. Yang, Z. Gan, L. Li, C. Li, X. Dai, H. Behl, J. Wang, L. Yuan, et al · 2023
Closest in time.