Fetching the paper…
Reading the bibliography…
We propose to compose dynamic tree structures that place the objects in an image into a visual context, helping visual reasoning tasks such as scene graph generation and visual Q&A.
Shortest connection networks and some generalizations
R. C. Prim · 1957
Earlier work this paper cites.
Scene perception: Detecting and judging objects undergoing relational violations
I. Biederman, R. J. Mezzanotte, and J. C. Rabinowitz · 1982
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Task-dependent influences of attention on the activation of human primary visual cortex
T. Watanabe, A. M. Harner, S. Miyauchi, Y. Sasaki, M. Nielsen, D. Palomo, and I. Mukai · 1998
Earlier work this paper cites.
Introduction to Algorithms
T. H. Cormen, C. Stein, R. L. Rivest, and C. E. Leiserson · 2001
Earlier work this paper cites.
Visual objects in context
M. Bar · 2004
Earlier work this paper cites.
The role of context in object recognition
A. Oliva and A. Torralba · 2007
Earlier work this paper cites.
An empirical study of context in object detection
S. K. Divvala, D. Hoiem, J. H. Hays, A. A. Efros, and M. Hebert · 2009
Earlier work this paper cites.
Beyond categories: The visual memex model for reasoning about object relationships
T. Malisiewicz and A. Efros · 2009
Earlier work this paper cites.
Parsing natural scenes and natural language with recursive neural networks
R. Socher, C. C. Lin, C. Manning, and A. Y. Ng · 2011
Earlier work this paper cites.
A tree-based context model for object recognition
M. J. Choi, A. Torralba, and A. S. Willsky · 2012
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, et al · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
K. S. Tai, R. Socher, and C. D. Manning · 2015
Earlier work this paper cites.
Saliency detection by multi-context deep learning
R. Zhao, W. Ouyang, H. Li, and X. Wang · 2015
Earlier work this paper cites.
Conditional random fields as recurrent neural networks
S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. Torr · 2015
Earlier work this paper cites.
Hadamard product for low-rank bilinear pooling
J.-H. Kim, K.-W. On, W. Lim, J. Kim, J.-W. Ha, and B.-T. Zhang · 2016
Cited alongside, same era.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Cited alongside, same era.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
Stacked hourglass networks for human pose estimation
A. Newell, K. Yang, and J. Deng · 2016
Cited alongside, same era.
Learning to refine object segments
P. O. Pinheiro, T.-Y. Lin, R. Collobert, and P. Dollár · 2016
Cited alongside, same era.
Multi-scale context aggregation by dilated convolutions
F. Yu and V. Koltun · 2016
Cited alongside, same era.
Scene graph generation by iterative message passing
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei · 2017
Later among the works it cites.
Relationship proposal networks
J. Zhang, M. Elhoseiny, S. Cohen, W. Chang, and A. M. Elgammal · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Closest in time.
Deep attention neural tensor network for visual question answering
Y. Bai, J. Fu, T. Zhao, and T. Mei · 2018
Closest in time.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2018
Closest in time.
Iterative visual reasoning beyond convolutions
X. Chen, L.-J. Li, L. Fei-Fei, and A. Gupta · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mutan: Multimodal tucker fusion for visual question answering
H. Ben-Younes, R. Cadene, M. Cord, and N. Thome · 2017
Cited alongside, same era.
Rethinking atrous convolution for semantic image segmentation
L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam · 2017
Cited alongside, same era.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Cited alongside, same era.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Cited alongside, same era.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Cited alongside, same era.
Closest in time.
Learning to compose task-specific tree structures
J. Choi, K. M. Yoo, and S.-g. Lee · 2018
Closest in time.
Tensorize, factorize and regularize: Robust visual relationship learning
S. Jae Hwang, S. N. Ravi, Z. Tao, H. J. Kim, M. D. Collins, and V. Singh · 2018
Closest in time.
A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár · 2018
Closest in time.
Factorizable net: An efficient subgraph-based framework for scene graph generation
Y. Li, W. Ouyang, B. Zhou, J. Shi, C. Zhang, and X. Wang · 2018
Closest in time.
Structure inference net: Object detection using scene-level context and instance-level relationships
Y. Liu, R. Wang, S. Shan, and X. Chen · 2018
Closest in time.
Question type guided attention in visual question answering
Y. Shi, T. Furlanello, S. Zha, and A. Anandkumar · 2018
Closest in time.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
D. Teney, P. Anderson, X. He, and A. van den Hengel · 2018
Closest in time.
Graph r-cnn for scene graph generation
J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh · 2018
Closest in time.
Exploring visual relationship for image captioning
T. Yao, Y. Pan, Y. Li, and T. Mei · 2018
Closest in time.
Zoom-net: Mining deep feature interactions for visual relationship recognition
G. Yin, L. Sheng, B. Liu, N. Yu, X. Wang, J. Shao, and C. C. Loy · 2018
Closest in time.
Neural motifs: Scene graph parsing with global context
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi · 2018
Closest in time.
Learning to count objects in natural images for visual question answering
Y. Zhang, J. Hare, and A. Prügel-Bennett · 2018
Closest in time.