Fetching the paper…
Reading the bibliography…
Spatial relationships between objects represent key scene information for humans to understand and interact with the world.
Indoor segmentation and support inference from rgbd images
Pushmeet Kohli Nathan Silberman, Derek Hoiem and Rob Fergus · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Detecting visual relationships with deep relational networks
Bo Dai, Yuqi Zhang, and Dahua Lin · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Vip-cnn: Visual phrase guided convolutional neural network
Yikang Li, Wanli Ouyang, Xiaogang Wang, and Xiao’ou Tang · 2017
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Christopher M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei · 2017
Earlier work this paper cites.
Position-aware attention and supervised data improve slot filling
Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning · 2017
Earlier work this paper cites.
Blender - a 3D modelling and rendering package
Blender Online Community · 2018
Earlier work this paper cites.
An intriguing failing of convolutional neural networks and the coordconv solution
Rosanne Liu, Joel Lehman, Piero Molino, Felipe Petroski Such, Eric Frank, Alex Sergeev, and Jason Yosinski · 2018
Earlier work this paper cites.
Neural motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2018
Earlier work this paper cites.
Link prediction based on graph neural networks
Muhan Zhang and Yixin Chen · 2018
Earlier work this paper cites.
Counterfactual critic multi-agent training for scene graph generation
Long Chen, Hanwang Zhang, Jun Xiao, Xiangnan He, Shiliang Pu, and Shih-Fu Chang · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Vrr-vg: Refocusing visually-relevant relationships
Yuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, and Tao Mei · 2019
Cited alongside, same era.
Relation prediction in knowledge graph by multi-label deep neural network
Yohei Onuki, Tsuyoshi Murata, Shun Nukui, Seiya Inagi, Xule Qiu, Masao Watanabe, and Hiroshi Okamoto · 2019
Cited alongside, same era.
Learning to compose dynamic tree structures for visual contexts
Kaihua Tang, Hanwang Zhang, Baoyuan Wu, Wenhan Luo, and Wei Liu · 2019
Cited alongside, same era.
Extracting multiple-relations in one-pass with pre-trained transformers
Haoyu Wang, Ming Tan, Mo Yu, Shiyu Chang, Dakuo Wang, Kun Xu, Xiaoxiao Guo, and Saloni Potdar · 2019
Cited alongside, same era.
Spatialsense: An adversarially crowdsourced benchmark for spatial relation recognition
Kaiyu Yang, Olga Russakovsky, and Jia Deng · 2019
Cited alongside, same era.
Shortcut learning in deep neural networks
Segmenter: Transformer for semantic segmentation
Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid · 2021
Later among the works it cites.
Visual distant supervision for scene graph generation
Yuan Yao, Ao Zhang, Xu Han, Mengdi Li, Cornelius Weber, Zhiyuan Liu, Stefan Wermter, and Maosong Sun · 2021
Later among the works it cites.
Document-level relation extraction with adaptive thresholding and localized context pooling
Wenxuan Zhou, Kevin Huang, Tengyu Ma, and Jing Huang · 2021
Later among the works it cites.
Learning perceptual concepts by bootstrapping from human queries
Andreea Bobu, Chris Paxton, Wei Yang, Balakumar Sundaralingam, Yu-Wei Chao, Maya Cakmak, and Dieter Fox · 2022
Later among the works it cites.
Resolving copycat problems in visual imitation learning via residual action prediction
Chia-Chi Chuang, Donglin Yang, Chuan Wen, and Yang Gao · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Cited alongside, same era.
Rel3d: A minimally contrastive benchmark for grounding spatial relations in 3d
Ankit Goyal, Kaiyu Yang, Dawei Yang, and Jia Deng · 2020
Cited alongside, same era.
The open images dataset v4
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al · 2020
Cited alongside, same era.
Gps-net: Graph property sensing network for scene graph generation
Xin Lin, Changxing Ding, Jinquan Zeng, and Dacheng Tao · 2020
Cited alongside, same era.
Reasoning with latent structure refinement for document-level relation extraction
Guoshun Nan, Zhijiang Guo, Ivan Sekulic, and Wei Lu · 2020
Cited alongside, same era.
Learning from Context or Names? An Empirical Study on Neural Relation Extraction
Hao Peng, Tianyu Gao, Xu Han, Yankai Lin, Peng Li, Zhiyuan Liu, Maosong Sun, and Jie Zhou · 2020
Cited alongside, same era.
Learning visual commonsense for robust scene graph generation
Alireza Zareian, Zhecan Wang, Haoxuan You, and Shih-Fu Chang · 2020
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Label semantic knowledge distillation for unbiased scene graph generation
Lin Li, Long Chen, Hanrong Shi, Wenxiao Wang, Jian Shao, Yi Yang, and Jun Xiao · 2022
Later among the works it cites.
Ru-net: regularized unrolling network for scene graph generation
Xin Lin, Changxing Ding, Jing Zhang, Yibing Zhan, and Dacheng Tao · 2022
Later among the works it cites.
Document-level relation extraction with adaptive focal loss and knowledge distillation
Qingyu Tan, Ruidan He, Lidong Bing, and Hwee Tou Ng · 2022
Later among the works it cites.
Generalizable task planning through representation pretraining
Chen Wang, Danfei Xu, and Li Fei-Fei · 2022
Later among the works it cites.
Fighting fire with fire: avoiding dnn shortcuts through priming
Chuan Wen, Jianing Qian, Jierui Lin, Jiaye Teng, Dinesh Jayaraman, and Yang Gao · 2022
Later among the works it cites.
Groupvit: Semantic segmentation emerges from text supervision
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang · 2022
Later among the works it cites.
Sornet: Spatial object-centric representations for sequential manipulation
Wentao Yuan, Chris Paxton, Karthik Desingh, and Dieter Fox · 2022
Later among the works it cites.
Image BERT pre-training with online tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechu Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Chahyon Ku, Carl Winge, Ryan Diaz, Wentao Yuan, and Karthik Desingh · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Genie: Generative interactive environments
Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al · 2024
Closest in time.