Fetching the paper…
Reading the bibliography…
People often refer to entities in an image in terms of their relationships with other entities.
Bidirectional recurrent neural networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
The Oxford handbook of compositionality
M. Werning, W. Hinzen, and E. Machery · 2012
Earlier work this paper cites.
Jointly learning to parse and perceive: Connecting natural language to the physical world
J. Krishnamurthy and T. Kollar · 2013
Earlier work this paper cites.
Selective search for object recognition
J. R. Uijlings, K. E. van de Sande, T. Gevers, and A. W. Smeulders · 2013
Earlier work this paper cites.
Fast and accurate shift-reduce constituent parsing
M. Zhu, Y. Zhang, W. Chen, M. Zhang, and J. Zhu · 2013
Earlier work this paper cites.
Multiscale combinatorial grouping
P. Arbeláez, J. Pont-Tuset, J. Barron, F. Marques, and J. Malik · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Geodesic object proposals
P. Krähenbühl and V. Koltun · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
C. D. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. J. Bethard, and D. McClosky · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Cited alongside, same era.
Edge boxes: Locating object proposals from edges
C. L. Zitnick and P. Dollár · 2014
Cited alongside, same era.
Multiple object recognition with visual attention
J. Ba, V. Mnih, and K. Kavukcuoglu · 2015
Cited alongside, same era.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Closest in time.
Segmentation from natural language expressions
R. Hu, M. Rohrbach, and T. Darrell · 2016
Closest in time.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Closest in time.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2016
Closest in time.
Ssd: Single shot multibox detector
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2016
Cited alongside, same era.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
R-fcn: Object detection via region-based fully convolutional networks
J. Dai, Y. Li, K. He, and J. Sun · 2016
Cited alongside, same era.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, and S. Reed · 2016
Closest in time.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Closest in time.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Closest in time.
Modeling context between objects for referring expression understanding
V. K. Nagaraja, V. I. Morariu, and L. S. Davis · 2016
Closest in time.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Closest in time.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Closest in time.