Fetching the paper…
Reading the bibliography…
Scene graph generation (SGG) aims to predict graph-structured descriptions of input images, in the form of objects and relationships between them.
An empirical study on leveraging scene graphs for visual question answering
C. Zhang, W.-L. Chao, and D. Xuan · 1907
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov · 2013
Earlier work this paper cites.
Attribute-based classification for zero-shot visual object categorization
C. H. Lampert, H. Nickisch, and S. Harmeling · 2013
Earlier work this paper cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. Van Der Maaten · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Earlier work this paper cites.
Latent embeddings for zero-shot classification
Y. Xian, Z. Akata, G. Sharma, Q. Nguyen, M. Hein, and B. Schiele · 2016
Earlier work this paper cites.
Joint embeddings of scene graphs and images
E. Belilovsky, M. Blaschko, J. Kiros, R. Urtasun, and R. Zemel · 2017
Earlier work this paper cites.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, and et al · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Cited alongside, same era.
Pixels to graphs by associative embedding
A. Newell and J. Deng · 2017
Cited alongside, same era.
Weakly-supervised learning of visual relations
J. Peyre, J. Sivic, I. Laptev, and C. Schmid · 2017
Cited alongside, same era.
P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Lio, and Y. Bengio · 2017
Cited alongside, same era.
Scene graph generation by iterative message passing
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei · 2017
Cited alongside, same era.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Cited alongside, same era.
Unpaired image captioning via scene graph alignments
J. Gu, S. Joty, J. Cai, H. Zhao, X. Yang, and G. Wang · 2019
Later among the works it cites.
Environmental drivers of systematicity and generalization in a situated agent, 2019
F. Hill, A. Lampinen, R. Schneider, S. Clark, M. Botvinick, J. L. McClelland, and A. Santoro · 2019
Later among the works it cites.
Understanding attention and generalization in graph neural networks
B. Knyazev, G. W. Taylor, and M. Amer · 2019
Later among the works it cites.
The truly deep graph convolutional networks for node classification
Y. Rong, W. Huang, T. Xu, and J. Huang · 2019
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blindfold baselines for embodied QA
A. Anand, E. Belilovsky, K. Kastner, H. Larochelle, and A. Courville · 2018
Cited alongside, same era.
Systematic generalization: what is required and can it be learned?
D. Bahdanau, S. Murty, M. Noukhovitch, T. H. Nguyen, H. de Vries, and A. Courville · 2018
Cited alongside, same era.
Scene graph generation via conditional random fields
W. Cong, W. Wang, and W.-C. Lee · 2018
Cited alongside, same era.
Learning conditioned graph structures for interpretable visual question answering
W. Norcliffe-Brown, S. Vafeias, and S. Parisot · 2018
Cited alongside, same era.
Neural motifs: Scene graph parsing with global context
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi · 2018
Cited alongside, same era.
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
C. Cangea, E. Belilovsky, P. Liò, and A. Courville · 2019
Cited alongside, same era.
Learning to compose dynamic tree structures for visual contexts
K. Tang, H. Zhang, B. Wu, W. Luo, and W. Liu · 2019
Later among the works it cites.
Probabilistic neural-symbolic models for interpretable visual question answering
R. Vedantam, K. Desai, S. Lee, M. Rohrbach, D. Batra, and D. Parikh · 2019
Later among the works it cites.
Generating expensive relationship features from cheap objects
X. Wang, Q. Sun, M. ANG, and T.-S. CHUA · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning
X. Yang, K. Tang, H. Zhang, and J. Cai · 2019
Later among the works it cites.
Unified vision-language pre-training for image captioning and VQA
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. J. Corso, and J. Gao · 2019
Later among the works it cites.
Scene graph benchmark in pytorch, 2020
K. Tang · 2020
Closest in time.
Unbiased scene graph generation from biased training
K. Tang, Y. Niu, J. Huang, J. Shi, and H. Zhang · 2020
Closest in time.