Fetching the paper…
Reading the bibliography…
Recent years have witnessed rapid progress in detecting and recognizing individual object instances.
Observing human-object interactions: Using spatial and functional compatibility for recognition
Abhinav Gupta, Aniruddha Kembhavi, and Larry S Davis · 2009
Earlier work this paper cites.
Action recognition from a distributed representation of pose and appearance
Subhransu Maji, Lubomir Bourdev, and Jitendra Malik · 2011
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
HICO: A benchmark for recognizing human-object interactions in images
Yu-Wei Chao, Zhan Wang, Yugeng He, Jiaxuan Wang, and Jia Deng · 2015
Earlier work this paper cites.
P-CNN: Pose-based cnn features for action recognition
Guilhem Chéron, Ivan Laptev, and Cordelia Schmid · 2015
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Contextual action recognition with r* cnn
Georgia Gkioxari, Ross Girshick, and Jitendra Malik · 2015
Earlier work this paper cites.
Saurabh Gupta and Jitendra Malik · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Weakly supervised deep detection networks
Hakan Bilen and Andrea Vedaldi · 2016
Earlier work this paper cites.
R-FCN: Object detection via region-based fully convolutional networks
Jifeng Dai, Yi Li, Kaiming He, and Jian Sun · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Learning models for actions and person-object interactions with transfer to question answering
Arun Mallya and Svetlana Lazebnik · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2017
Weakly-supervised learning of visual relations
Julia Peyre, Ivan Laptev, Cordelia Schmid, and Josef Sivic · 2017
Later among the works it cites.
Phrase localization and visual relationship detection with comprehensive linguistic cues
Bryan A Plummer, Arun Mallya, Christopher M Cervantes, Julia Hockenmaier, and Svetlana Lazebnik · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Scene graph generation by iterative message passing
Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei · 2017
Later among the works it cites.
PPR-FCN: Weakly supervised visual relation detection via parallel pairwise r-fcn
Hanwang Zhang, Zawlin Kyaw, Jinyang Yu, and Shih-Fu Chang · 2017
Later among the works it cites.
Towards context-aware interaction recognition for visual relationship detection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning to detect human-object interactions
Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng · 2017
Cited alongside, same era.
DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille · 2017
Cited alongside, same era.
Detecting visual relationships with deep relational networks
Bo Dai, Yuqi Zhang, and Dahua Lin · 2017
Cited alongside, same era.
Attentional pooling for action recognition
Rohit Girdhar and Deva Ramanan · 2017
Cited alongside, same era.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks
Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko · 2017
Cited alongside, same era.
Bohan Zhuang, Lingqiao Liu, Chunhua Shen, and Ian Reid · 2017
Later among the works it cites.
Detectron
Ross Girshick, Ilija Radosavovic, Georgia Gkioxari, Piotr Dollár, and Kaiming He · 2018
Closest in time.
Detecting and recognizing human-object interactions
Georgia Gkioxari, Ross Girshick, Piotr Dollár, and Kaiming He · 2018
Closest in time.
Learn to pay attention
Saumya Jetley, Nicholas A Lord, Namhoon Lee, and Philip HS Torr · 2018
Closest in time.
Detecting visual relationships using box attention
Alexander Kolesnikov, Christoph H Lampert, and Vittorio Ferrari · 2018
Closest in time.
Scaling human-object interaction recognition through zero-shot learning
Liyue Shen, Serena Yeung, Judy Hoffman, Greg Mori, and Li Fei-Fei · 2018
Closest in time.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Closest in time.
Neural Motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2018
Closest in time.