Fetching the paper…
Reading the bibliography…
We present a simple method that achieves unexpectedly superior performance for Complex Reasoning involved Visual Question Answering.
Towards context-aware interaction recognition for visual relationship detection
B. Zhuang, L. Liu, C. Shen, and I. Reid · 1907
Earlier work this paper cites.
GloVe: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Faster r-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Earlier work this paper cites.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Earlier work this paper cites.
Vip-cnn: A visual phrase reasoning convolutional neural network for visual relationship detection
Y. Li, W. Ouyang, and X. Wang · 2017
Cited alongside, same era.
Scene graph generation by iterative message passing
D. Xu, Y. Zhu, C. Choy, and L. Fei-Fei · 2017
Cited alongside, same era.
Visual relationship detection with internal and external linguistic knowledge distillation
R. Yu, A. Li, V. I. Morariu, and L. S. Davis · 2017
Cited alongside, same era.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Cited alongside, same era.
Relationship proposal networks
J. Zhang, M. Elhoseiny, S. Cohen, W. Chang, and A. Elgammal · 2017
Cited alongside, same era.
Compositional attention networks for machine reasoning
D. A. Hudson and C. D. Manning · 2018
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
K. Yi, J. Wu, C. Gan, A. Torralba, P. Kohli, and J. Tenenbaum · 2018
Later among the works it cites.
Zoom-net: Mining deep feature interactions for visual relationship recognition
G. Yin, L. Sheng, B. Liu, N. Yu, X. Wang, J. Shao, and C. Change Loy · 2018
Later among the works it cites.
Neural motifs: Scene graph parsing with global context
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi · 2018
Later among the works it cites.
An interpretable model for scene graph generation
J. Zhang, K. Shih, A. Tao, B. Catanzaro, and A. Elgammal · 2018
Later among the works it cites.
Introduction to the 1st place winning model of openimages relationship detection challenge
J. Zhang, K. Shih, A. Tao, B. Catanzaro, and A. Elgammal · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
TVQA: localized, compositional video question answering
J. Lei, L. Yu, M. Bansal, and T. L. Berg · 2018
Cited alongside, same era.
Large-scale visual relationship understanding
J. Zhang, Y. Kalantidis, M. Rohrbach, M. Paluri, A. Elgammal, and M. Elhoseiny · 2019
Closest in time.
Graphical contrastive losses for scene graph generation
J. Zhang, K. J. Shih, A. Elgammal, A. Tao, and B. Catanzaro · 2019
Closest in time.