Fetching the paper…
Reading the bibliography…
Panoptic Scene Graph Generation (PSG) parses objects and predicts their relationships (predicate) to connect human language and visual scenes.
VQA: Visual Question Answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Image Retrieval Using Scene Graphs
Johnson, J.; Krishna, R.; Stark, M.; Li, L.-J.; Shamma, D.; Bernstein, M.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2017 · 2017
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Graph-Structured Representations for Visual Question Answering
Teney, D.; Liu, L.; and van den Hengel, A. 2017 · 2017
Earlier work this paper cites.
Scene Graph Generation by Iterative Message Passing
Xu, D.; Zhu, Y.; Choy, C. B.; and Fei-Fei, L. 2017 · 2017
Earlier work this paper cites.
Neural Motifs: Scene Graph Parsing With Global Context
Zellers, R.; Yatskar, M.; Thomson, S.; and Choi, Y. 2018 · 2018
Earlier work this paper cites.
Knowledge-Embedded Routing Network for Scene Graph Generation
Chen, T.; Yu, W.; Chen, R.; and Lin, L. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Panoptic Segmentation
Kirillov, A.; He, K.; Girshick, R.; Rother, C.; and Dollar, P. 2019 · 2019
Cited alongside, same era.
Learning to Compose Dynamic Tree Structures for Visual Contexts
Tang, K.; Zhang, H.; Wu, B.; Luo, W.; and Liu, W. 2019 · 2019
Cited alongside, same era.
Invariant Risk Minimization
Arjovsky, M.; Bottou, L.; Gulrajani, I.; and Lopez-Paz, D. 2020 · 2020
Cited alongside, same era.
GPS-Net: Graph Property Sensing Network for Scene Graph Generation
Lin, X.; Ding, C.; Zeng, J.; and Tao, D. 2020 · 2020
Cited alongside, same era.
Unbiased Scene Graph Generation From Biased Training
Tang, K.; Niu, Y.; Huang, J.; Shi, J.; and Zhang, H. 2020 · 2020
Cited alongside, same era.
CogTree: Cognition Tree Loss for Unbiased Scene Graph Generation
Yu, J.; Chai, Y.; Hu, Y.; and Wu, Q. 2020 · 2020
Cited alongside, same era.
Video as Conditional Graph Hierarchy for Multi-Granular Question Answering
Xiao, J.; Yao, A.; Liu, Z.; Li, Y.; Ji, W.; and Chua, T.-S. 2022 · 2022
Later among the works it cites.
Panoptic Scene Graph Generation
Yang, J.; Ang, Y. Z.; Guo, Z.; Zhou, K.; Zhang, W.; and Liu, Z. 2022 · 2022
Later among the works it cites.
PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models
Yao, Y.; Chen, Q.; Zhang, A.; Ji, W.; Liu, Z.; Chua, T.-S.; and Sun, M. 2022 · 2022
Later among the works it cites.
Fine-Grained Scene Graph Generation with Data Transfer
Zhang, A.; Yao, Y.; Chen, Q.; Ji, W.; Liu, Z.; Sun, M.; and Chua, T.-S. 2022 · 2022
Later among the works it cites.
Causal Property based Anti-Conflict Modeling with Hybrid Data Augmentation for Unbiased Scene Graph Generation
Zhang, R.; and An, G. 2022 · 2022
Later among the works it cites.
A Comprehensive Survey of Scene Graphs: Generation and Application
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
Yao, Y.; Zhang, A.; Zhang, Z.; Liu, Z.; Chua, T.-S.; and Sun, M. 2021 · 2021
Cited alongside, same era.
Multi-modal cross-domain alignment network for video moment retrieval
Fang, X.; Liu, D.; Zhou, P.; and Hu, Y. 2022 · 2022
Cited alongside, same era.
Personalizing Intervened Network for Long-tailed Sequential User Behavior Modeling
Lv, Z.; Wang, F.; Zhang, S.; Kuang, K.; Yang, H.; and Wu, F. 2022 · 2022
Cited alongside, same era.
Rethinking the Two-Stage Framework for Grounded Situation Recognition
Wei, M.; Chen, L.; Ji, W.; Yue, X.; and Chua, T.-S. 2022 · 2022
Cited alongside, same era.
Say As You Wish: Fine-Grained Control of Image Caption Generation With Abstract Scene Graphs
Chen, S.; Jin, Q.; Wang, P.; and Wu, Q. 2020a
Cited in the paper.
A Simple Framework for Contrastive Learning of Visual Representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020b
Cited in the paper.
Chang, X.; Ren, P.; Xu, P.; Li, Z.; Chen, X.; and Hauptmann, A. 2023 · 2023
Closest in time.
You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos
Fang, X.; Liu, D.; Zhou, P.; and Nan, G. 2023 · 2023
Closest in time.
DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model Generalization
Lv, Z.; Zhang, W.; Zhang, S.; Kuang, K.; Wang, F.; Wang, Y.; Chen, Z.; Shen, T.; Yang, H.; Ooi, B. C.; et al. 2023 · 2023
Closest in time.
Revisiting the Domain Shift and Sample Uncertainty in Multi-source Active Domain Transfer
Zhang, W.; Lv, Z.; Zhou, H.; Liu, J.-W.; Li, J.; Li, M.; Tang, S.; and Zhuang, Y. 2023 · 2023
Closest in time.