Fetching the paper…
Reading the bibliography…
This paper revisits human-object interaction (HOI) recognition at image level without using supervisions of object location and human pose.
A framework for multiple-instance learning
O. Maron and T. Lozano-Pérez · 1998
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Contextual action recognition with r* cnn
G. Gkioxari, R. Girshick, and J. Malik · 2015
Earlier work this paper cites.
Hico: A benchmark for recognizing human-object interactions in images
Y.-W. Chao, Z. Wang, Y. He, J. Wang, and J. Deng · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Learning models for actions and person-object interactions with transfer to question answering
A. Mallya and S. Lazebnik · 2016
Earlier work this paper cites.
Scene graph generation by iterative message passing
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei · 2017
Earlier work this paper cites.
Attentional pooling for action recognition
R. Girdhar and D. Ramanan · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Normface: L2 hypersphere embedding for face verification
F. Wang, X. Xiang, J. Cheng, and A. L. Yuille · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Pairwise body-part attention for recognizing human-object interactions
H.-S. Fang, J. Cao, Y.-W. Tai, and C. Lu · 2018
Cited alongside, same era.
Learning to detect human-object interactions
Y.-W. Chao, Y. Liu, X. Liu, H. Zeng, and J. Deng · 2018
Cited alongside, same era.
ican: Instance-centric attention network for human-object interaction detection
C. Gao, Y. Zou, and J.-B. Huang · 2018
Cited alongside, same era.
Detecting and recognizing human-object interactions
G. Gkioxari, R. Girshick, P. Dollár, and K. He · 2018
Cited alongside, same era.
Factorizable net: an efficient subgraph-based framework for scene graph generation
Y. Li, W. Ouyang, B. Zhou, J. Shi, C. Zhang, and X. Wang · 2018
Cited alongside, same era.
Graph r-cnn for scene graph generation
J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh · 2018
Amplifying key cues for human-object-interaction detection
Y. Liu, Q. Chen, and A. Zisserman · 2020
Later among the works it cites.
Spatially conditioned graphs for detecting human-object interactions
F. Z. Zhang, D. Campbell, and S. Gould · 2020
Later among the works it cites.
Ppdm: Parallel point detection and matching for real-time human-object interaction detection
Y. Liao, S. Liu, F. Wang, Y. Chen, C. Qian, and J. Feng · 2020
Later among the works it cites.
Uniondet: Union-level detector towards real-time human-object interaction detection
B. Kim, T. Choi, J. Kang, and H. J. Kim · 2020
Later among the works it cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural motifs: Scene graph parsing with global context
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Hake: Human activity knowledge engine
Y.-L. Li, L. Xu, X. Liu, X. Huang, Y. Xu, M. Chen, Z. Ma, S. Wang, H.-S. Fang, and C. Lu · 2019
Cited alongside, same era.
Knowledge-embedded routing network for scene graph generation
T. Chen, W. Yu, R. Chen, and L. Lin · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
K. Tang, Y. Niu, J. Huang, J. Shi, and H. Zhang · 2020
Later among the works it cites.
Vivo: Surpassing human performance in novel object captioning with visual vocabulary pre-training
X. Hu, X. Yin, K. Lin, L. Wang, L. Zhang, J. Gao, and Z. Liu · 2020
Later among the works it cites.
Circle loss: A unified perspective of pair similarity optimization
Y. Sun, C. Cheng, Y. Zhang, C. Zhang, L. Zheng, Z. Wang, and Y. Wei · 2020
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Closest in time.
Glance and gaze: Inferring action-aware points for one-stage human-object interaction detection
X. Zhong, X. Qu, C. Ding, and D. Tao · 2021
Closest in time.
Reformulating hoi detection as adaptive set prediction
M. Chen, Y. Liao, S. Liu, Z. Chen, F. Wang, and C. Qian · 2021
Closest in time.
End-to-end human object interaction detection with hoi transformer
C. Zou, B. Wang, Y. Hu, J. Liu, Q. Wu, Y. Zhao, B. Li, C. Zhang, C. Zhang, Y. Wei, and J. Sun · 2021
Closest in time.
Qpic: Query-based pairwise human-object interaction detection with image-wide contextual information
M. Tamura, H. Ohashi, and T. Yoshinaga · 2021
Closest in time.