Fetching the paper…
Reading the bibliography…
Panoptic Scene Graph Generation (PSG) aims at achieving a comprehensive image understanding by simultaneously segmenting objects and predicting relations among objects.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Detecting visual relationships with deep relational networks
Bo Dai, Yuqi Zhang, and Dahua Lin · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Weakly-supervised learning of visual relations
Julia Peyre, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei · 2017
Earlier work this paper cites.
Visual translation embedding network for visual relation detection
Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, and Tat-Seng Chua · 2017
Earlier work this paper cites.
Image understanding using vision and reasoning through scene description graph
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos, and Cornelia Fermüller · 2018
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
Holger Caesar, Jasper Uijlings, and Vittorio Ferrari · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Image captioning with scene-graph based semantic concepts
Lizhao Gao, Bo Wang, and Wenmin Wang · 2018
Earlier work this paper cites.
Tensorize, factorize and regularize: Robust visual relationship learning
Seong Jae Hwang, Sathya N Ravi, Zirui Tao, Hyunwoo J Kim, Maxwell D Collins, and Vikas Singh · 2018
Earlier work this paper cites.
Neural motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2018
Earlier work this paper cites.
Counterfactual critic multi-agent training for scene graph generation
Long Chen, Hanwang Zhang, Jun Xiao, Xiangnan He, Shiliang Pu, and Shih-Fu Chang · 2019
Earlier work this paper cites.
Scene graph generation with external knowledge and image reconstruction
Jiuxiang Gu, Handong Zhao, Zhe Lin, Sheng Li, Jianfei Cai, and Mingyang Ling · 2019
Earlier work this paper cites.
Neural message passing for visual relationship detection
Yue Hu, Siheng Chen, Xu Chen, Ya Zhang, and Xiao Gu · 2019
Earlier work this paper cites.
Panoptic segmentation
Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollár · 2019
Earlier work this paper cites.
Natural language guided visual relationship detection
Wentong Liao, Bodo Rosenhahn, Ling Shuai, and Michael Ying Yang · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Attentive relational networks for mapping images to scene graphs
Mengshi Qi, Weijian Li, Zhengyuan Yang, Yunhong Wang, and Jiebo Luo · 2019
Cited alongside, same era.
Explainable and explicit visual reasoning over scene graphs
Jiaxin Shi, Hanwang Zhang, and Juanzi Li · 2019
Cited alongside, same era.
Learning to compose dynamic tree structures for visual contexts
Kaihua Tang, Hanwang Zhang, Baoyuan Wu, Wenhan Luo, and Wei Liu · 2019
Cited alongside, same era.
Visual relation detection with multi-level attention
Sipeng Zheng, Shizhe Chen, and Qin Jin · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Say as you wish: Fine-grained control of image caption generation with abstract scene graphs
Shizhe Chen, Qin Jin, Peng Wang, and Qi Wu · 2020
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Later among the works it cites.
Panoptic scene graph generation
Jingkang Yang, Yi Zhe Ang, Zujin Guo, Kaiyang Zhou, Wayne Zhang, and Ziwei Liu · 2022
Later among the works it cites.
Fine-grained scene graph generation with data transfer
Ao Zhang, Yuan Yao, Qianyu Chen, Wei Ji, Zhiyuan Liu, Maosong Sun, and Tat-Seng Chua · 2022
Later among the works it cites.
Claude, 2023
Anthropic · 2023
Closest in time.
Large language models are visual reasoning coordinators
Liangyu Chen, Bo Li, Sheng Shen, Jingkang Yang, Chunyuan Li, Kurt Keutzer, Trevor Darrell, and Ziwei Liu · 2023
Closest in time.
Reltr: Relation transformer for scene graph generation
Yuren Cong, Michael Ying Yang, and Bodo Rosenhahn · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visual relationship detection with low rank non-negative tensor decomposition
Mohammed Haroon Dupty, Zhen Zhang, and Wee Sun Lee · 2020
Cited alongside, same era.
Scene graph reasoning for visual question answering
Marcel Hildebrandt, Hang Li, Rajat Koner, Volker Tresp, and Stephan Günnemann · 2020
Cited alongside, same era.
Contextual translation embedding for visual relationship detection and scene graph generation
Zih-Siou Hung, Arun Mallya, and Svetlana Lazebnik · 2020
Cited alongside, same era.
Gps-net: Graph property sensing network for scene graph generation
Xin Lin, Changxing Ding, Jinquan Zeng, and Dacheng Tao · 2020
Cited alongside, same era.
Unbiased scene graph generation from biased training
Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang · 2020
Cited alongside, same era.
Bridging knowledge graphs to generate scene graphs
Alireza Zareian, Svebor Karaman, and Shih-Fu Chang · 2020
Cited alongside, same era.
Closest in time.
Bard, 2023
Google · 2023
Closest in time.
Inject semantic concepts into image tagging for open-set recognition
Xinyu Huang, Yi-Jie Huang, Youcai Zhang, Weiwei Tian, Rui Feng, Yuejie Zhang, Yanchun Xie, Yaqian Li, and Lei Zhang · 2023
Closest in time.
Skew class-balanced re-weighting for unbiased scene graph generation
Haeyong Kang and Chang D Yoo · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Closest in time.
Lisa: Reasoning segmentation via large language model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia · 2023
Closest in time.
Panoptic scene graph generation with semantics-prototype learning
Li Li, Wei Ji, Yiming Wu, Mengze Li, You Qin, Lina Wei, and Roger Zimmermann · 2023
Closest in time.
What large language models bring to text-rich vqa?
Xuejing Liu, Wei Tang, Xinzhe Ni, Jinghui Lu, Rui Zhao, Zechao Li, and Fei Tan · 2023
Closest in time.
Recent advances in natural language processing via large pre-trained language models: A survey
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth · 2023
Closest in time.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Dhruv Shah, Błażej Osiński, Sergey Levine, et al · 2023
Closest in time.
Scene graph contrastive learning for embodied navigation
Kunal Pratap Singh, Jordi Salvador, Luca Weihs, and Aniruddha Kembhavi · 2023
Closest in time.
Cotdet: Affordance knowledge prompting for task driven object detection
Jiajin Tang, Ge Zheng, Jingyi Yu, and Sibei Yang · 2023
Closest in time.
Multimodal large language model for visual navigation
Yao-Hung Hubert Tsai, Vansh Dhar, Jialu Li, Bowen Zhang, and Jian Zhang · 2023
Closest in time.
A simple baseline for knowledge-based visual question answering
Alexandros Xenos, Themos Stafylakis, Ioannis Patras, and Georgios Tzimiropoulos · 2023
Closest in time.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao · 2023
Closest in time.
Visually-prompted language model for fine-grained scene graph generation in an open world
Qifan Yu, Juncheng Li, Yu Wu, Siliang Tang, Wei Ji, and Yueting Zhuang · 2023
Closest in time.
Next-chat: An lmm for chat, detection and segmentation
Ao Zhang, Liming Zhao, Chen-Wei Xie, Yun Zheng, Wei Ji, and Tat-Seng Chua · 2023
Closest in time.