Fetching the paper…
Reading the bibliography…
Grounding referring expressions in images aims to locate the object instance in an image described by a referring expression.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,” IEEE Transactions on Signal Processing , vol. 45, no. 11, pp. 2673–2681, 1997
1997
Earlier work this paper cites.
R. Cadene, H. Ben-Younes, M. Cord, and N. Thome, “Murel: Multimodal relational reasoning for visual question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 1989–1998
1998
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. A. Forsyth, “Every picture tells a story: generating sentences from images,” in European conference on computer vision , 2010, pp. 15–29
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg, “Babytalk: Understanding and generating simple image descriptions,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 35, no. 12, pp. 2891–2903, Dec. 2013. [Online]. Available: doi.ieeecomputersociety.org/10.1109/TPAMI.2012.162
2012
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 787–798
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European conference on computer vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
C. D. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. J. Bethard, and D. McClosky, “The Stanford CoreNLP natural language processing toolkit,” in Association for Computational Linguistics (ACL) System Demonstrations , 2014, pp. 55–60
2014
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Advances in neural information processing systems , 2015, pp. 91–99
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” The IEEE International Conference on Computer Vision (ICCV) , pp. 2425–2433, 2015
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” 2015, pp. 3156–3164
2015
Earlier work this paper cites.
M. Malinowski, M. Rohrbach, and M. Fritz, “Ask your neurons: A neural-based approach to answering questions about images,” The IEEE International Conference on Computer Vision (ICCV) , pp. 1–9, 2015
2015
Earlier work this paper cites.
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu, “Are you talking to a machine? dataset and methods for multilingual image question answering,” Advances in neural information processing systems , pp. 2296–2304, 2015
2015
Earlier work this paper cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , 2015, pp. 815–823
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , 2016, pp. 11–20
2016
Earlier work this paper cites.
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele, “Grounding of textual phrases in images by reconstruction,” in European Conference on Computer Vision . Springer, 2016, pp. 817–834
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in European Conference on Computer Vision . Springer, 2016, pp. 69–85
2016
Earlier work this paper cites.
V. K. Nagaraja, V. I. Morariu, and L. S. Davis, “Modeling context between objects for referring expression understanding,” in European Conference on Computer Vision . Springer, 2016, pp. 792–807
2016
Earlier work this paper cites.
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei, “Visual relationship detection with language priors,” in European Conference on Computer Vision . Springer, 2016, pp. 852–869
2016
Earlier work this paper cites.
S. Bell, C. Lawrence Zitnick, K. Bala, and R. Girshick, “Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 2874–2883
2016
Earlier work this paper cites.
G. Li and Y. Yu, “Visual saliency detection based on multiscale deep cnn features,” IEEE Transactions on Image Processing , vol. 25, no. 11, pp. 5012–5024, 2016
2016
Earlier work this paper cites.
A. Jain, A. R. Zamir, S. Savarese, and A. Saxena, “Structural-rnn: Deep learning on spatio-temporal graphs,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 5308–5317
2016
Earlier work this paper cites.
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach, “Multimodal compact bilinear pooling for visual question answering and visual grounding,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , 2016, pp. 457–468
2016
Cited alongside, same era.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” Advances in neural information processing systems , pp. 289–297, 2016
2016
Cited alongside, same era.
K. J. Shih, S. Singh, and D. Hoiem, “Where to look: Focus regions for visual question answering,” IEEE conference on computer vision and pattern recognition (CVPR) , pp. 4613–4621, 2016
2016
Cited alongside, same era.
Z. Yang, X. He, J. Gao, L. Deng, and A. J. Smola, “Stacked attention networks for image question answering,” IEEE conference on computer vision and pattern recognition (CVPR) , pp. 21–29, 2016
2016
Cited alongside, same era.
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg, “Mattnet: Modular attention network for referring expression comprehension,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
H. Zhang, Y. Niu, and S.-F. Chang, “Grounding referring expressions in images by variational context,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 4158–4166
2018
Later among the works it cites.
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi, “Neural motifs: Scene graph parsing with global context,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 5831–5840
2018
Later among the works it cites.
C. Deng, Q. Wu, Q. Wu, F. Hu, F. Lyu, and M. Tan, “Visual grounding via accumulated attention,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein, “Neural module networks,” IEEE conference on computer vision and pattern recognition (CVPR) , pp. 39–48, 2016
2016
Cited alongside, same era.
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei, “Visual7w: Grounded question answering in images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4995–5004
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , 2016, pp. 770–778
2016
Cited alongside, same era.
2016
Cited alongside, same era.
J. Liu, L. Wang, and M.-H. Yang, “Referring expression generation and comprehension via attributes,” in The IEEE International Conference on Computer Vision (ICCV) , Oct 2017
2017
Cited alongside, same era.
L. Yu, H. Tan, M. Bansal, and T. L. Berg, “A joint speakerlistener-reinforcer model for referring expressions,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , vol. 2, 2017
2017
Cited alongside, same era.
R. Luo and G. Shakhnarovich, “Comprehension-guided referring expressions,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , vol. 2, 2017
2017
Cited alongside, same era.
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko, “Modeling relationships in referential expressions with compositional modular networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) . IEEE, 2017, pp. 4418–4427
2017
Cited alongside, same era.
L. Chen, G. Papandreou, I. Kokkinos, K. P. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, no. 4, pp. 834–848, 2018
2018
Later among the works it cites.
H. Zhang, K. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal, “Context encoding for semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
G. Li and Y. Yu, “Contrast-oriented deep neural networks for salient object detection,” IEEE Transactions on Neural Networks and Learning Systems , vol. 29, no. 12, pp. 6038–6051, 2018
2018
Later among the works it cites.
Y. Liu, R. Wang, S. Shan, and X. Chen, “Structure inference net: Object detection using scene-level context and instance-level relationships,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 6985–6994
2018
Later among the works it cites.
X. Wu, G. Li, Q. Cao, Q. Ji, and L. Lin, “Interpretable video captioning via trajectory structured localization,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , 2018, pp. 6829–6837
2018
Later among the works it cites.
B. Dai, S. Fidler, and D. Lin, “A neural compositional paradigm for image captioning,” 2018
2018
Later among the works it cites.
C. Wu, J. Liu, X. Wang, and X. Dong, “Chain of reasoning for visual question answering,” in Advances in Neural Information Processing Systems 31 , S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds. Curran Associates, Inc., 2018, pp. 273–283
2018
Later among the works it cites.
Q. Cao, X. Liang, B. Li, G. Li, and L. Lin, “Visual question reasoning on general dependency tree,” IEEE conference on computer vision and pattern recognition (CVPR) , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Lio, and Y. Bengio, “Graph attention networks,” International Conference on Learning Representations, 2018 , 2018
2018
Later among the works it cites.
X. Wang, Y. Ye, and A. Gupta, “Zero-shot recognition via semantic embeddings and knowledge graphs,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6857–6866
2018
Later among the works it cites.
W. Norcliffe-Brown, S. Vafeias, and S. Parisot, “Learning conditioned graph structures for interpretable visual question answering,” in Advances in Neural Information Processing Systems , 2018, pp. 8334–8343
2018
Later among the works it cites.
J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh, “Graph r-cnn for scene graph generation,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 670–685
2018
Later among the works it cites.
T. Yao, Y. Pan, Y. Li, and T. Mei, “Exploring visual relationship for image captioning,” 2018
2018
Later among the works it cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , 2018
2018
Later among the works it cites.
S. Yang, G. Li, and Y. Yu, “Cross-modal relationship inference for grounding referring expressions,” in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR) , 2019, pp. 4145–4154
2019
Closest in time.
G. Li, Y. Gan, H. Wu, N. Xiao, and L. Lin, “Cross-modal attentional context learning for rgb-d object detection,” IEEE Transactions on Image Processing , vol. 28, no. 4, pp. 1591–1601, 2019
2019
Closest in time.
Y. Zhang, J. C. Niebles, and A. Soto, “Interpretable visual question answering by visual grounding from attention supervision mining,” in IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, 2019, pp. 349–357
2019
Closest in time.
X. Yang, K. Tang, H. Zhang, and J. Cai, “Auto-encoding scene graphs for image captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 10 685–10 694
2019
Closest in time.
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” 2015, pp. 2048–2057
2057
Closest in time.