Fetching the paper…
Reading the bibliography…
Referring object detection and referring image segmentation are important tasks that require joint understanding of visual information and natural language.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and F. Li · 2009
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Earlier work this paper cites.
VQA: visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, T. Darrell, and K. Saenko · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and F. Li · 2015
Earlier work this paper cites.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Segmentation from natural language expressions
R. Hu, M. Rohrbach, and T. Darrell · 2016
Earlier work this paper cites.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Cited alongside, same era.
Question relevance in VQA: identifying non-visual and false-premise questions
A. Ray, G. Christie, M. Bansal, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Cited alongside, same era.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. S. Bernstein, and L. Fei-Fei · 2016
Comprehension-guided referring expressions
R. Luo and G. Shakhnarovich · 2017
Later among the works it cites.
A simple neural network module for relational reasoning
A. Santoro, D. Raposo, D. G. T. Barrett, M. Malinowski, R. Pascanu, P. Battaglia, and T. Lillicrap · 2017
Later among the works it cites.
A joint speaker-listener-reinforcer model for referring expressions
L. Yu, H. Tan, M. Bansal, and T. L. Berg · 2017
Later among the works it cites.
Visual referring expression recognition: What do systems actually learn?
V. Cirik, L. Morency, and T. Berg-Kirkpatrick · 2018
Later among the works it cites.
Explainable neural computation via stack neural module networks
R. Hu, J. Andreas, T. Darrell, and K. Saenko · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2017
Cited alongside, same era.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. B. Girshick · 2017
Cited alongside, same era.
Inferring and executing programs for visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, L. Fei-Fei, C. L. Zitnick, and R. B. Girshick · 2017
Cited alongside, same era.
Recurrent multimodal interaction for referring image segmentation
C. Liu, Z. Lin, X. Shen, J. Yang, X. Lu, and A. L. Yuille · 2017
Cited alongside, same era.
D. A. Hudson and C. D. Manning · 2018
Later among the works it cites.
Referring image segmentation via recurrent refinement networks
R. Li, K. Li, Y.-C. Kuo, M. Shu, X. Qi, X. Shen, and J. Jia · 2018
Later among the works it cites.
Dynamic multimodal instance segmentation guided by natural language queries
E. Margffoy-Tuay, J. C. Pérez, E. Botero, and P. Arbeláez · 2018
Later among the works it cites.
Transparency by design: Closing the gap between performance and interpretability in visual reasoning
D. Mascharka, P. Tran, R. Soklaski, and A. Majumdar · 2018
Later among the works it cites.
Learning by asking questions
I. Misra, R. B. Girshick, R. Fergus, M. Hebert, A. Gupta, and L. van der Maaten · 2018
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. C. Courville · 2018
Later among the works it cites.
Mattnet: Modular attention network for referring expression comprehension
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg · 2018
Later among the works it cites.