Fetching the paper…
Reading the bibliography…
Supervised or weakly supervised methods for phrase localization (textual grounding) either rely on human annotations or some other supervised models, e.g., object detectors.
G. A. Miller, “WordNet: a lexical database for English,” Communications of the ACM
1995
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Proc. CVPR
2009
Earlier work this paper cites.
C. H. Lampert, M. B. Blaschko, and T. Hofmann, “Efficient Subwindow Search: A Branch and Bound Framework for Object Localization,” PAMI
2009
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The PASCAL visual object classes (VOC) challenge,” IJCV
2010
Earlier work this paper cites.
R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, and S. Süsstrunk, “Slic superpixels compared to state-of-the-art superpixel methods,” IEEE TPAMI
2012
Earlier work this paper cites.
J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders, “Selective search for object recognition,” IJCV
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Proc. ECCV
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” TACL
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg, “ReferItGame: Referring to objects in photographs of natural scenes,” in Proc. EMNLP
2014
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in Proc. ICCV
2015
Earlier work this paper cites.
L. Wang, Y. Li, and S. Lazebnik, “Learning deep structure-preserving image-text embeddings,” in CVPR
2016
Earlier work this paper cites.
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell, “Natural language object retrieval,” in Proc. CVPR
2016
Earlier work this paper cites.
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach, “Multimodal compact bilinear pooling for visual question answering and visual grounding,” in EMNLP
2016
Earlier work this paper cites.
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele, “Grounding of textual phrases in images by reconstruction,” in ECCV
2016
Earlier work this paper cites.
B. A. Plummer, A. Mallya, C. M. Cervantes, J. Hockenmaier, and S. Lazebnik, “Phrase localization and visual relationship detection with comprehensive image-language cues,” in Proc. ICCV
2017
Cited alongside, same era.
R. A. Yeh, J. Xiong, W.-M. Hwu, M. Do, and A. G. Schwing, “Interpretable and globally optimal prediction for textual grounding using image concepts,” in Proc. NeurIPS
2017
Cited alongside, same era.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al
2017
Cited alongside, same era.
2017
Cited alongside, same era.
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko, “Modeling relationships in referential expressions with compositional modular networks,” in Proc. CVPR
A. Sadhu, K. Chen, and R. Nevatia, “Zero-shot grounding of objects from natural language queries,” in Proc. ICCV
2019
Later among the works it cites.
Z. Yang, T. Chen, L. Wang, and J. Luo, “Improving one-stage visual grounding by recursive sub-query construction,” in Proc. ECCV
2020
Later among the works it cites.
L. Parcalabescu and A. Frank, “Exploring phrase grounding without training: Contextualisation and extension to text-based image retrieval,” in Proc. CVPRW
2020
Later among the works it cites.
Z. Mu, S. Tang, J. Tan, Q. Yu, and Y. Zhuang, “Disentangled motif-aware graph learning for phrase grounding,” in Proc. AAAI
2021
Later among the works it cites.
L. Wang, J. Huang, Y. Li, K. Xu, Z. Yang, and D. Yu, “Improving weakly supervised visual grounding by contrastive knowledge distillation,” in Proc. CVPR
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
K. Endo, M. Aono, E. Nichols, and K. Funakoshi, “An attention-based regression model for grounding textual phrases in images,” in Proc. IJCAI
2017
Cited alongside, same era.
∗ equal contribution
K. Chen ∗ , R. Kovvuri ∗ , and R. Nevatia, “Query-guided regression network with context policy for phrase grounding,” in Proc. ICCV · 2017
Cited alongside, same era.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” TPAMI
2017
Cited alongside, same era.
L. Wang, Y. Li, J. Huang, and S. Lazebnik, “Learning two-branch neural networks for image-text matching tasks,” IEEE TPAMI
2018
Cited alongside, same era.
B. A. Plummer, P. Kordas, M. H. Kiapour, S. Zheng, R. Piramuthu, and S. Lazebnik, “Conditional image-text embedding networks,” in Proc. ECCV
2018
Cited alongside, same era.
R. A. Yeh, M. N. Do, and A. G. Schwing, “Unsupervised textual grounding: Linking words to image concepts,” in Proc. CVPR
2018
Cited alongside, same era.
S. Yang, G. Li, and Y. Yu, “Dynamic graph attention for referring expression comprehension,” in Proc. ICCV
2019
Cited alongside, same era.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in Proc. ICML
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui, “Zero-shot detection via vision and language knowledge distillation,” arXiv e-prints
2021
Later among the works it cites.
A. Zareian, K. D. Rosa, D. H. Hu, and S.-F. Chang, “Open-vocabulary object detection using captions,” in Proc. CVPR
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.