Fetching the paper…
Reading the bibliography…
Most existing work that grounds natural language phrases in images starts with the assumption that the phrase in question is relevant to the image.
H. Hotelling, “Relations between two sets of variates,” Biometrika , pp. 321–377, 1936
1936
Earlier work this paper cites.
P. J. Huber, “Robust estimation of a location parameter,” Annals of Statistics , vol. 53, no. 1, pp. 73–101, 1964
1964
Earlier work this paper cites.
G. A. Miller, “Wordnet: A lexical database for english,” Communications of the ACM , vol. 38, no. 11, pp. 39–41, 1995
1995
Earlier work this paper cites.
M. Grubinger, P. Clough, H. Müller, and T. Deselaers, “The IAPR TC-12 benchmark – a new evaluation resource for visual information systems,” 2006
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR , 2009
2009
Earlier work this paper cites.
2013
Earlier work this paper cites.
G. Andrew, R. Arora, J. Bilmes, and K. Livescu, “Deep canonical correlation analysis,” in ICML , 2013
2013
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in EMNLP , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in ECCV , 2014
2014
Earlier work this paper cites.
Y. Gong, Q. Ke, M. Isard, and S. Lazebnik, “A multi-view embedding space for modeling internet images, tags, and their semantics,” IJCV , vol. 106, no. 2, pp. 210–233, 2014
2014
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” TACL , vol. 2, pp. 67–78, 2014
2014
Earlier work this paper cites.
P. Bojanowski, R. Lajugie, E. Grave, F. Bach, I. Laptev, J. Ponce, and C. Schmid, “Weakly-Supervised Alignment of Video With Text,” in ICCV , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in NIPS , 2015
2015
Earlier work this paper cites.
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollar, J. Gao, X. He, M. Mitchell, J. Platt, L. Zitnick, and G. Zweig, “From captions to visual concepts and back,” in CVPR , 2015
2015
Earlier work this paper cites.
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei, “Image retrieval using scene graphs,” in CVPR , 2015
2015
Earlier work this paper cites.
R. Girshick, “Fast R-CNN,” in ICCV , 2015
2015
Earlier work this paper cites.
B. Klein, G. Lev, G. Sadeh, and L. Wolf, “Associating neural word embeddings with deep image representations using fisher vector,” in CVPR , 2015
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015
2015
Cited alongside, same era.
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in ICML , 2015
2015
Cited alongside, same era.
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach, “Multimodal compact bilinear pooling for visual question answering and visual grounding,” in EMNLP , 2016
2016
Cited alongside, same era.
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell, “Natural language object retrieval,” in CVPR , 2016
R. A. Yeh, J. Xiong, W. mei W. Hwu, M. N. Do, and A. G. Schwing, “Interpretable and globally optimal prediction for textual grounding using image concepts,” in NIPS , 2017
2017
Later among the works it cites.
Y. Zhang, L. Yuan, Y. Guo, Z. He, I.-A. Huang, and H. Lee, “Discriminative bimodal networks for visual localization and detection with natural language queries,” in CVPR , 2017
2017
Later among the works it cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei, “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV , 2017
2017
Later among the works it cites.
R. Luo and G. Shakhnarovich, “Comprehension-guided referring expressions,” in CVPR , 2017
2017
Later among the works it cites.
J. Liu, L. Wang, and M.-H. Yang, “Referring expression generation and comprehension via attributes,” in ICCV , 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in CVPR , 2016
2016
Cited alongside, same era.
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele, “Grounding of textual phrases in images by reconstruction,” in ECCV , 2016
2016
Cited alongside, same era.
M. Wang, M. Azab, N. Kojima, R. Mihalcea, and J. Deng, “Structured matching for phrase localization,” in ECCV , 2016
2016
Cited alongside, same era.
L. Wang, Y. Li, and S. Lazebnik, “Learning deep structure-preserving image-text embeddings,” in CVPR , 2016
2016
Cited alongside, same era.
I. Misra, C. L. Zitnick, M. Mitchell, and R. Girshick, “Seeing through the Human Reporting Bias: Visual Classifiers from Noisy Human-Centric Labels,” in CVPR , 2016
2016
Cited alongside, same era.
J. Johnson, A. Karpathy, and L. Fei-Fei, “Densecap: Fully convolutional localization networks for dense captioning,” in CVPR , 2016
2016
Cited alongside, same era.
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei, “Visual relationship detection with language priors,” in ECCV , 2016
2016
Cited alongside, same era.
2017
Later among the works it cites.
T. Yao, Y. Pan, Y. Li, Z. Qiu, and T. Mei, “Boosting image captioning with attributes,” in ICCV , 2017
2017
Later among the works it cites.
C. Liu, J. Mao, F. Sha, and A. Yuille, “Attention correctness in neural image captioning,” in AAAI , 2017
2017
Later among the works it cites.
B. A. Plummer, P. Kordas, M. H. Kiapour, S. Zheng, R. Piramuthu, and S. Lazebnik, “Conditional image-text embedding networks,” in ECCV , 2018
2018
Closest in time.
R. Hinami and S. Satoh, “Discriminative learning of open-vocabulary object retrieval and localization by negative phrase augmentation,” in EMNLP , 2018
2018
Closest in time.
2018
Closest in time.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR , 2018
2018
Closest in time.
K.-H. Lee, X. Chen, G. Hua, H. Hu, and X. He, “Stacked cross attention for image-text matching,” in ECCV , 2018
2018
Closest in time.
J. Wehrmann and R. C. Barros, “Bidirectional retrieval made simple,” in CVPR , 2018
2018
Closest in time.
A. Gupta, P. Dollar, and R. Girshick, “LVIS: A dataset for large vocabulary instance segmentation,” in CVPR , 2019, pp. 5356–5364
2019
Closest in time.
M. Acharya, K. Jariwala, and C. Kanan, “VQD: Visual query detection in natural scenes,” in NAACL , 2019
2019
Closest in time.
A. Burns, R. Tan, K. Saenko, S. Sclaroff, and B. A. Plummer, “Language features matter: Effective language representations for vision-language tasks,” in ICCV , 2019
2019
Closest in time.
M. Bajaj, L. Wang, and L. Sigal, “G3raphground: Graph-based language grounding,” in ICCV , 2019
2019
Closest in time.
Z. Yang, B. Gong, L. Wang, W. Huang, D. Yu, and J. Luo, “A fast and accurate one-stage approach to visual grounding,” in ICCV , 2019
2019
Closest in time.