Fetching the paper…
Reading the bibliography…
The extraction of a scene graph with objects as nodes and mutual relationships as edges is the basis for a deep understanding of image content.
In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pp. 1532–1543 (2014)
Pennington, J., Socher, R., Manning, C.D.: Glove: Global vectors for word representation · 2014
Earlier work this paper cites.
arXiv preprint arXiv:1409.1556 (2014)
Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition · 2014
Earlier work this paper cites.
The journal of machine learning research 15
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., Salakhutdinov, R.: Dropout: a simple way to prevent neural networks from overfitting · 2014
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3668–3678 (2015)
Johnson, J., Krishna, R., Stark, M., Li, L.J., Shamma, D., Bernstein, M., Fei-Fei, L.: Image retrieval using scene graphs · 2015
Earlier work this paper cites.
In: Proceedings of the IEEE international conference on computer vision, pp. 2641–2649 (2015)
Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., Lazebnik, S.: Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models · 2015
Earlier work this paper cites.
In: Advances in neural information processing systems, pp. 91–99 (2015)
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks · 2015
Earlier work this paper cites.
arXiv preprint arXiv:1505.00853 (2015)
Xu, B., Wang, N., Chen, T., Li, M.: Empirical evaluation of rectified activations in convolutional network · 2015
Earlier work this paper cites.
arXiv preprint arXiv:1607.06450 (2016)
Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization · 2016
Earlier work this paper cites.
In: European conference on computer vision, pp. 852–869. Springer (2016)
Lu, C., Krishna, R., Bernstein, M., Fei-Fei, L.: Visual relationship detection with language priors · 2016
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 11–20 (2016)
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A.L., Murphy, K.: Generation and comprehension of unambiguous object descriptions · 2016
Earlier work this paper cites.
In: European Conference on Computer Vision, pp. 792–807. Springer (2016)
Nagaraja, V.K., Morariu, V.I., Davis, L.S.: Modeling context between objects for referring expression understanding · 2016
Earlier work this paper cites.
In: International Semantic Web Conference, pp. 53–68. Springer (2017)
Baier, S., Ma, Y., Tresp, V.: Improving visual relationship detection using semantic modeling of scene descriptions · 2017
Earlier work this paper cites.
In: Proceedings of the IEEE international conference on computer vision, pp. 2961–2969 (2017)
He, K., Gkioxari, G., Dollár, P., Girshick, R.: Mask r-cnn · 2017
Earlier work this paper cites.
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1115–1124 (2017)
Hu, R., Rohrbach, M., Andreas, J., Darrell, T., Saenko, K.: Modeling relationships in referential expressions with compositional modular networks · 2017
Earlier work this paper cites.
International Journal of Computer Vision 123
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al.: Visual genome: Connecting language and vision using crowdsourced dense image annotations · 2017
Earlier work this paper cites.
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1347–1356 (2017)
Li, Y., Ouyang, W., Wang, X., Tang, X.: Vip-cnn: Visual phrase guided convolutional neural network · 2017
Earlier work this paper cites.
In: Advances in neural information processing systems, pp. 2171–2180 (2017)
Newell, A., Deng, J.: Pixels to graphs by associative embedding · 2017
Earlier work this paper cites.
In: Advances in neural information processing systems, pp. 5998–6008 (2017)
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need · 2017
Earlier work this paper cites.
In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2017)
Xu, D., Zhu, Y., Choy, C.B., Fei-Fei, L.: Scene graph generation by iterative message passing · 2017
Cited alongside, same era.
In: Proceedings of the IEEE international conference on computer vision, pp. 1974–1982 (2017)
Yu, R., Li, A., Morariu, V.I., Davis, L.S.: Visual relationship detection with internal and external linguistic knowledge distillation · 2017
Cited alongside, same era.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5532–5540 (2017)
Zhang, H., Kyaw, Z., Chang, S.F., Chua, T.S.: Visual translation embedding network for visual relation detection · 2017
Cited alongside, same era.
In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-up and top-down attention for image captioning and visual question answering · 2018
Cited alongside, same era.
In: Advances in Neural Information Processing Systems, pp. 13–23 (2019)
Lu, J., Batra, D., Parikh, D., Lee, S.: Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks · 2019
Later among the works it cites.
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3957–3966 (2019)
Qi, M., Li, W., Yang, Z., Wang, Y., Luo, J.: Attentive relational networks for mapping images to scene graphs · 2019
Later among the works it cites.
Sharifzadeh, S., Baharlou, S.M., Berrendorf, M., Koner, R., Tresp, V.: Improving visual relation detection using depth maps (2019)
2019
Later among the works it cites.
In: Conference on Computer Vision and Pattern Recognition (2019)
Tang, K., Zhang, H., Wu, B., Luo, W., Liu, W.: Learning to compose dynamic tree structures for visual contexts · 2019
Later among the works it cites.
arXiv preprint arXiv:1905.09418 (2019)
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., Titov, I.: Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding · 2018
Cited alongside, same era.
In: Advances in Neural Information Processing Systems, pp. 7211–7221 (2018)
Herzig, R., Raboh, M., Chechik, G., Berant, J., Globerson, A.: Mapping images to scene graphs with permutation-invariant structured prediction · 2018
Cited alongside, same era.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6985–6994 (2018)
Liu, Y., Wang, R., Shan, S., Chen, X.: Structure inference net: Object detection using scene-level context and instance-level relationships · 2018
Cited alongside, same era.
In: IJCAI, pp. 949–956 (2018)
Wan, H., Luo, Y., Peng, B., Zheng, W.S.: Representation learning for scene graph completion via jointly structural and visual embedding · 2018
Cited alongside, same era.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7794–7803 (2018)
Wang, X., Girshick, R., Gupta, A., He, K.: Non-local neural networks · 2018
Cited alongside, same era.
In: Advances in Neural Information Processing Systems, pp. 560–570 (2018)
Woo, S., Kim, D., Cho, D., Kweon, I.S.: Linknet: Relational embedding for scene graph · 2018
Cited alongside, same era.
In: Proceedings of the European conference on computer vision (ECCV), pp. 670–685 (2018)
Yang, J., Lu, J., Lee, S., Batra, D., Parikh, D.: Graph r-cnn for scene graph generation · 2018
Cited alongside, same era.
In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 322–338 (2018)
Yin, G., Sheng, L., Liu, B., Yu, N., Wang, X., Shao, J., Change Loy, C.: Zoom-net: Mining deep feature interactions for visual relationship recognition · 2018
Cited alongside, same era.
Later among the works it cites.
Journal of Visual Communication and Image Representation 58
Xu, N., Liu, A.A., Liu, J., Nie, W., Su, Y.: Scene graph captioner: Image captioning based on structural visual representation · 2019
Later among the works it cites.
In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, pp. 9185–9194 (2019)
Zhang, J., Kalantidis, Y., Rohrbach, M., Paluri, M., Elgammal, A., Elhoseiny, M.: Large-scale visual relationship understanding · 2019
Later among the works it cites.
Zhang, J., Shih, K.J., Elgammal, A., Tao, A., Catanzaro, B.: Graphical contrastive losses for scene graph parsing · 2019
Later among the works it cites.
arXiv preprint arXiv:2005.12872 (2020)
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers · 2020
Closest in time.
arXiv preprint arXiv:2007.01072 (2020)
Hildebrandt, M., Li, H., Koner, R., Tresp, V., Günnemann, S.: Scene graph reasoning for visual question answering · 2020
Closest in time.
arXiv preprint arXiv:2005.08230 (2020)
Knyazev, B., de Vries, H., Cangea, C., Taylor, G.W., Courville, A., Belilovsky, E.: Graph density-aware losses for novel compositions in scene graph generation · 2020
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3746–3753 (2020)
Lin, X., Ding, C., Zeng, J., Tao, D.: Gps-net: Graph property sensing network for scene graph generation · 2020
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 178–179 (2020)
Schroeder, B., Tripathi, S.: Structured query-based image retrieval using scene graphs · 2020
Closest in time.
In: Conference on Computer Vision and Pattern Recognition (2020)
Tang, K., Niu, Y., Huang, J., Shi, J., Zhang, H.: Unbiased scene graph generation from biased training · 2020
Closest in time.
arXiv preprint arXiv:2006.09623 (2020)
Zareian, A., Wang, Z., You, H., Chang, S.F.: Learning visual commonsense for robust scene graph generation · 2020
Closest in time.
Koner, R., Li, H., Hildebrandt, M., Das, D., Tresp, V., Günnemann, S.: Graphhopper: Multi-hop scene graph reasoning for visual question answering (2021)
2021
Closest in time.
Koner, R., Sinhamahapatra, P., Roscher, K., Günnemann, S., Tresp, V.: Oodformer: Out-of-distribution detection transformer (2021)
2021
Closest in time.
arXiv preprint arXiv:2107.05448 (2021)
Koner, R., Sinhamahapatra, P., Tresp, V.: Scenes and surroundings: Scene graph generation using relation transformer · 2021
Closest in time.