Fetching the paper…
Reading the bibliography…
Understanding the spatial relations between objects in images is a surprisingly challenging task.
Whence and whither in spatial language and spatial cognition?
Barbara Landau and Ray Jackendoff · 1993
Earlier work this paper cites.
Active learning with gaussian processes for object categorization
Ashish Kapoor, Kristen Grauman, Raquel Urtasun, and Trevor Darrell · 2007
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Beat the machine: Challenging workers to find the unknown unknowns
Josh Attenberg, Panagiotis G Ipeirotis, and Foster J Provost · 2011
Earlier work this paper cites.
Recognition using visual phrases
Mohammad Amin Sadeghi and Ali Farhadi · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Grounding spatial relations for human-robot interaction
Sergio Guadarrama, Lorenzo Riano, Dave Golland, Daniel Go, Yangqing Jia, Dan Klein, Pieter Abbeel, Trevor Darrell, et al · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Large-scale live active learning: Training object detectors with crawled data and crowds
Sudheendra Vijayanarasimhan and Kristen Grauman · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Stacked hourglass networks for human pose estimation
Alejandro Newell, Kaiyu Yang, and Jia Deng · 2016
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Wei-Lun Chao, Hexiang Hu, and Fei Sha · 2017
Cited alongside, same era.
Detecting visual relationships with deep relational networks
Bo Dai, Yuqi Zhang, and Dahua Lin · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Weakly-supervised learning of visual relations
Julia Peyre, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2017
Later among the works it cites.
Scene graph generation by iterative message passing
Danfei Xu, Yuke Zhu, Christopher B. Choy, and Li Fei-Fei · 2017
Later among the works it cites.
Visual relationship detection with internal and external linguistic knowledge distillation
Ruichi Yu, Ang Li, Vlad I. Morariu, and Larry S. Davis · 2017
Later among the works it cites.
Visual translation embedding network for visual relation detection
Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, and Tat-Seng Chua · 2017
Later among the works it cites.
Ppr-fcn: Weakly supervised visual relation detection via parallel pairwise r-fcn
Hanwang Zhang, Zawlin Kyaw, Jinyang Yu, and Shih-Fu Chang · 2017
Later among the works it cites.
Towards context-aware interaction recognition for visual relationship detection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
An analysis of visual question answering algorithms
Kushal Kafle and Christopher Kanan · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2017
Cited alongside, same era.
Vip-cnn: Visual phrase guided convolutional neural network
Yikang Li, Wanli Ouyang, Xiaogang Wang, and Xiao’ou Tang · 2017
Cited alongside, same era.
Scene graph generation from objects, phrases and region captions
Yikang Li, Wanli Ouyang, Bolei Zhou, Kun Wang, and Xiaogang Wang · 2017
Cited alongside, same era.
Deep variation-structured reinforcement learning for visual relationship and attribute detection
Xiaodan Liang, Lisa Lee, and Eric P. Xing · 2017
Cited alongside, same era.
Bohan Zhuang, Lingqiao Liu, Chunhua Shen, and Ian Reid · 2017
Later among the works it cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Later among the works it cites.
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Tom Duerig, and Vittorio Ferrari · 2018
Later among the works it cites.
Factorizable net: an efficient subgraph-based framework for scene graph generation
Yikang Li, Wanli Ouyang, Bolei Zhou, Jianping Shi, Chao Zhang, and Xiaogang Wang · 2018
Later among the works it cites.
Linknet: Relational embedding for scene graph
Sanghyun Woo, Dahun Kim, Donghyeon Cho, and In So Kweon · 2018
Later among the works it cites.
Shuffle-then-assemble: Learning object-agnostic visual relationship features
Xu Yang, Hanwang Zhang, and Jianfei Cai · 2018
Later among the works it cites.
Semantic robot programming for goal-directed manipulation in cluttered scenes
Zhen Zeng, Zheming Zhou, Zhiqiang Sui, and Odest Chadwicke Jenkins · 2018
Later among the works it cites.
Explicit bias discovery in visual question answering models
Varun Manjunatha, Nirat Saini, and Larry S Davis · 2019
Closest in time.