Fetching the paper…
Reading the bibliography…
Scene Graph Generation (SGG) aims to structurally and comprehensively represent objects and their connections in images, it can significantly benefit scene understanding and other related downstream tasks.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of Machine Learning Research , 2008
2008
Earlier work this paper cites.
M. E. ElAlami, “A novel image retrieval model based on the most relevant features,” Knowledge-Based Systems , vol. 24, no. 1, pp. 23–32, 2011
2011
Earlier work this paper cites.
D. Rafailidis, S. Manolopoulou, and P. Daras, “A unified framework for multimodal retrieval,” Pattern Recognition , vol. 46, no. 12, pp. 3358–3370, 2013
2013
Earlier work this paper cites.
M. Turk, “Multimodal interaction: A review,” Pattern recognition letters , vol. 36, pp. 189–195, 2014
2014
Earlier work this paper cites.
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei, “Image retrieval using scene graphs,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 3668–3678
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” Advances in Neural Information Processing Systems , pp. 91–99, 2015
2015
Earlier work this paper cites.
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei, “Visual relationship detection with language priors,” in Proceedings of the European Conference on Computer Vision , 2016, pp. 852–869
2016
Earlier work this paper cites.
M. Long, Y. Cao, J. Wang, and P. S. Yu, “Composite correlation quantization for efficient multimodal retrieval,” in Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval , 2016, pp. 579–588
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Teney, L. Liu, and A. van Den Hengel, “Graph-structured representations for visual question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 1–9
2017
Earlier work this paper cites.
B. Dai, Y. Zhang, and D. Lin, “Detecting visual relationships with deep relational networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 3298–3308
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International Journal of Computer Vision , pp. 32–73, 2017
2017
Earlier work this paper cites.
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua, “Visual translation embedding network for visual relation detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 5532–5540
2017
Earlier work this paper cites.
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei, “Scene graph generation by iterative message passing,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 5410–5419
2017
Earlier work this paper cites.
D. Ramachandram and G. W. Taylor, “Deep multimodal learning: A survey on recent advances and trends,” IEEE signal processing magazine , vol. 34, no. 6, pp. 96–108, 2017
2017
Earlier work this paper cites.
X. Ochoa, A. C. Lang, and G. Siemens, “Multimodal learning analytics,” The handbook of learning analytics , vol. 1, pp. 129–141, 2017
2017
Earlier work this paper cites.
L. Chen, S. Srivastava, Z. Duan, and C. Xu, “Deep cross-modal audio-visual generation,” in Proceedings of the on Thematic Workshops of ACM Multimedia 2017 , 2017, pp. 349–357
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Soleymani, D. Garcia, B. Jou, B. Schuller, S.-F. Chang, and M. Pantic, “A survey of multimodal sentiment analysis,” Image and Vision Computing , vol. 65, pp. 3–14, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , pp. 5998–6008, 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 2117–2125
2017
Earlier work this paper cites.
J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh, “Graph r-cnn for scene graph generation,” in Proceedings of the European Conference on Computer Vision , 2018, pp. 670–685
2018
Cited alongside, same era.
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi, “Neural motifs: Scene graph parsing with global context,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 5831–5840
2018
Cited alongside, same era.
Y. Xian, C. H. Lampert, B. Schiele, and Z. Akata, “Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 9, pp. 2251–2265, 2018
2018
Cited alongside, same era.
X. Yang, H. Zhang, and J. Cai, “Shuffle-then-assemble: Learning object-agnostic visual relationship features,” in Proceedings of the European conference on computer vision , 2018, pp. 36–52
2018
Cited alongside, same era.
S. Yan, C. Shen, Z. Jin, J. Huang, R. Jiang, Y. Chen, and X.-S. Hua, “Pcpl: Predicate-correlation perception learning for unbiased scene graph generation,” in Proceedings of the ACM International Conference on Multimedia , 2020, pp. 265–273
2020
Later among the works it cites.
S. Narayan, A. Gupta, F. S. Khan, C. G. Snoek, and L. Shao, “Latent embedding feedback and discriminative features for zero-shot classification,” in European Conference on Computer Vision . Springer, 2020, pp. 479–495
2020
Later among the works it cites.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 736–740
2020
Later among the works it cites.
V. Gabeur, C. Sun, K. Alahari, and C. Schmid, “Multi-modal transformer for video retrieval,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16 . Springer, 2020, pp. 214–229
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
B. Wang, L. Ma, W. Zhang, and W. Liu, “Reconstruction network for video captioning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7622–7631
2018
Cited alongside, same era.
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He, “Attngan: Fine-grained text to image generation with attentional generative adversarial networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 1316–1324
2018
Cited alongside, same era.
Y. Li, M. Min, D. Shen, D. Carlson, and L. Carin, “Video generation from text,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Cited alongside, same era.
C. Deng, Q. Wu, Q. Wu, F. Hu, F. Lyu, and M. Tan, “Visual grounding via accumulated attention,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 7746–7755
2018
Cited alongside, same era.
2018
Cited alongside, same era.
K. Tang, H. Zhang, B. Wu, W. Luo, and W. Liu, “Learning to compose dynamic tree structures for visual contexts,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 6619–6628
2019
Cited alongside, same era.
T. Chen, W. Yu, R. Chen, and L. Lin, “Knowledge-embedded routing network for scene graph generation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 6163–6171
2019
Cited alongside, same era.
2020
Later among the works it cites.
X. Lin, C. Ding, J. Zeng, and D. Tao, “Gps-net: Graph property sensing network for scene graph generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 3746–3753
2020
Later among the works it cites.
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov et al. , “The open images dataset v4,” International Journal of Computer Vision , vol. 128, no. 7, pp. 1956–1981, 2020
2020
Later among the works it cites.
A. Hogan, E. Blomqvist, M. Cochez, C. d’Amato, G. d. Melo, C. Gutierrez, S. Kirrane, J. E. L. Gayo, R. Navigli, S. Neumaier et al. , “Knowledge graphs,” ACM Computing Surveys (CSUR) , vol. 54, no. 4, pp. 1–37, 2021
2021
Later among the works it cites.
Y. Luo, J. Ji, X. Sun, L. Cao, Y. Wu, F. Huang, C.-W. Lin, and R. Ji, “Dual-level collaborative transformer for image captioning,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2021, pp. 2286–2293
2021
Later among the works it cites.
R. Li, S. Zhang, B. Wan, and X. He, “Bipartite graph network with adaptive message passing for unbiased scene graph generation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 11 109–11 119
2021
Later among the works it cites.
M. Suhail, A. Mittal, B. Siddiquie, C. Broaddus, J. Eledath, G. Medioni, and L. Sigal, “Energy-based learning for scene graph generation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 13 936–13 945
2021
Later among the works it cites.
Y. Guo, L. Gao, X. Wang, Y. Hu, X. Xu, X. Lu, H. T. Shen, and J. Song, “From general to specific: Informative scene graph generation via balance adjustment,” in Proceedings of the IEEE International Conference on Computer Vision. , 2021, pp. 16 383–16 392
2021
Later among the works it cites.
B. Knyazev, H. de Vries, C. Cangea, G. W. Taylor, A. Courville, and E. Belilovsky, “Generative compositional augmentations for scene graph prediction,” in Proceedings of the IEEE International Conference on Computer Vision , 2021, pp. 15 827–15 837
2021
Later among the works it cites.
H. Liu, N. Yan, M. Mortazavi, and B. Bhanu, “Fully convolutional scene graph generation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 11 546–11 556
2021
Later among the works it cites.
G. Yang, J. Zhang, Y. Zhang, B. Wu, and Y. Yang, “Probabilistic modeling of semantic ambiguity for scene graph generation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2021, pp. 12 527–12 536
2021
Later among the works it cites.
J. Wang, Z. Zeng, B. Chen, T. Dai, and S.-T. Xia, “Contrastive quantization with code memory for unsupervised image retrieval,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2022, pp. 2468–2476
2022
Later among the works it cites.
Z. Fei, “Attention-aligned transformer for image captioning,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2022, pp. 607–615
2022
Later among the works it cites.
A. Cherian, C. Hori, T. K. Marks, and J. Le Roux, “(2.5+ 1) d spatio-temporal scene graphs for video question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2022, pp. 444–453
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Goel, B. Fernando, F. Keller, and H. Bilen, “Not all relations are equal: Mining informative labels for scene graph generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 15 596–15 606
2022
Later among the works it cites.
W. Li, H. Zhang, Q. Bai, G. Zhao, N. Jiang, and X. Yuan, “Ppdl: Predicate probability distribution based loss for unbiased scene graph generation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2022, pp. 19 447–19 456
2022
Later among the works it cites.