Fetching the paper…
Reading the bibliography…
We investigate the problem of producing structured graph representations of visual scenes.
Edge and curve detection for visual scene analysis
A. Rosenfeld and M. Thurston · 1971
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
What, where and who? classifying events by scene and object recognition
L.-J. Li et al · 2007
Earlier work this paper cites.
Objects in context
A. Rabinovich, A. Vedaldi, C. Galleguillos, E. Wiewiora, and S. Belongie · 2007
Earlier work this paper cites.
An empirical study of context in object detection
S. K. Divvala, D. Hoiem, J. H. Hays, A. A. Efros, and M. Hebert · 2009
Earlier work this paper cites.
Actions in context
M. Marszalek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Context based object categorization: A critical survey
C. Galleguillos and S. Belongie · 2010
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
B. Yao et al · 2010
Earlier work this paper cites.
Approaching the symbol grounding problem with probabilistic graphical models
S. Tellex, T. Kollar, S. Dickerson, M. R. Walter, A. G. Banerjee, S. Teller, and N. Roy · 2011
Earlier work this paper cites.
Neil: Extracting visual knowledge from web data
X. Chen, A. Shrivastava, and A. Gupta · 2013
Earlier work this paper cites.
A survey on still image based human action recognition
G. Guo et al · 2014
Earlier work this paper cites.
What are you talking about? text-to-image coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Reasoning about object affordances in a knowledge base representation
Y. Zhu, A. Fathi, and L. Fei-Fei · 2014
Earlier work this paper cites.
VQA: visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Hico: A benchmark for recognizing human-object interactions in images
Y.-W. Chao, Z. Wang, Y. He, J. Wang, and J. Deng · 2015
Earlier work this paper cites.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. Lawrence Zitnick · 2015
Earlier work this paper cites.
Building a large-scale multimodal knowledge base for visual question answering
Z. et al · 2015
Earlier work this paper cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollar, J. Gao, X. He, M. Mitchell, J. C. Platt, C. Lawrence Zitnick, and G. Zweig · 2015
Cited alongside, same era.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. e. a. Gao · 2015
Cited alongside, same era.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. Plummer, L. Wang, C. Cervantes, J. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
Annotating object instances with a polygon-rnn
L. Castrejon, K. Kundu, R. Urtasun, and S. Fidler · 2017
Closest in time.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Closest in time.
Deep semantic role labeling: What works and what’s next
L. He, K. Lee, M. Lewis, and L. Zettlemoyer · 2017
Closest in time.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Closest in time.
Hadamard Product for Low-rank Bilinear Pooling
J.-H. Kim, K. W. On, W. Lim, J. Kim, J.-W. Ha, and B.-T. Zhang · 2017
Closest in time.
Neural amr: Sequence-to-sequence models for parsing and generation
I. Konstas, S. Iyer, M. Yatskar, Y. Choi, and L. Zettlemoyer · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Image question answering: A visual semantic embedding model and a new dataset
M. Ren et al · 2015
Cited alongside, same era.
Viske: Visual knowledge extraction and question answering by visual verification of relation phrases
F. Sadeghi, S. K. Divvala, and A. Farhadi · 2015
Cited alongside, same era.
Training very deep networks
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Cited alongside, same era.
Grammar as a foreign language
O. Vinyals, Ł. Kaiser, T. Koo, S. Petrov, I. Sutskever, and G. Hinton · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Visual madlibs: Fill in the blank image generation and question answering
L. e. a. Yu · 2015
Cited alongside, same era.
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Closest in time.
Vip-cnn: Visual phrase guided convolutional neural network
Y. Li, W. Ouyang, X. Wang, et al · 2017
Closest in time.
ViP-CNN: Visual Phrase Guided Convolutional Neural Network
Y. Li, W. Ouyang, X. Wang, and X. Tang · 2017
Closest in time.
Scene graph generation from objects, phrases and region captions
Y. Li, W. Ouyang, B. Zhou, K. Wang, and X. Wang · 2017
Closest in time.
Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
X. Liang, L. Lee, and E. P. Xing · 2017
Closest in time.
Pixels to graphs by associative embedding
A. Newell and J. Deng · 2017
Closest in time.
Graph-structured representations for visual question answering
D. Teney, L. Liu, and A. van den Hengel · 2017
Closest in time.
Scene Graph Generation by Iterative Message Passing
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei · 2017
Closest in time.
Obj2text: Generating visually descriptive language from object layouts
X. Yin and V. Ordonez · 2017
Closest in time.
Obj2text: Generating visually descriptive language from object layouts
X. Yin and V. Ordonez · 2017
Closest in time.
Visual relationship detection with internal and external linguistic knowledge distillation
R. Yu, A. Li, V. I. Morariu, and L. S. Davis · 2017
Closest in time.
Zero-shot activity recognition with verb attribute induction
R. Zellers and Y. Choi · 2017
Closest in time.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Closest in time.