Fetching the paper…
Reading the bibliography…
We have witnessed significant progress in human-object interaction (HOI) detection.
Diagnosing error in object detectors
Derek Hoiem, Yodsawalai Chodpathumwan, and Qieyun Dai · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
HICO: A benchmark for recognizing human-object interactions in images
Yu-Wei Chao, Zhan Wang, Yugeng He, Jiaxuan Wang, and Jia Deng · 2015
Earlier work this paper cites.
Saurabh Gupta and Jitendra Malik · 2015
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Where to look: Focus regions for visual question answering
Kevin J Shih, Saurabh Singh, and Derek Hoiem · 2016
Earlier work this paper cites.
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Fvqa: Fact-based visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Earlier work this paper cites.
Convolutional image captioning
Jyoti Aneja, Aditya Deshpande, and Alexander G Schwing · 2018
Earlier work this paper cites.
Learning to detect human-object interactions
Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng · 2018
Earlier work this paper cites.
Fine-tuning cnn image retrieval with no human annotation
Filip Radenović, Giorgos Tolias, and Ondřej Chum · 2018
Earlier work this paper cites.
Unsupervised image captioning
Yang Feng, Lin Ma, Wei Liu, and Jiebo Luo · 2019
Earlier work this paper cites.
No-frills human-object interaction detection: Factorization, layout encodings, and training techniques
Tanmay Gupta, Alexander Schwing, and Derek Hoiem · 2019
Cited alongside, same era.
Entangled transformer for image captioning
Guang Li, Linchao Zhu, Ping Liu, and Yi Yang · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun · 2019
Cited alongside, same era.
Detect-to-retrieve: Efficient regional aggregation for image search
Marvin Teichmann, Andre Araujo, Menglong Zhu, and Jack Sim · 2019
Cited alongside, same era.
Relation parsing neural network for human-object interaction detection
Penghao Zhou and Mingmin Chi · 2019
Cited alongside, same era.
TIDE: A general toolbox for identifying object detection errors
Detecting human-object interaction via fabricated compositional learning
Zhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng, and Dacheng Tao · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Qpic: Query-based pairwise human-object interaction detection with image-wide contextual information
Masato Tamura, Hiroki Ohashi, and Tomoaki Yoshinaga · 2021
Later among the works it cites.
End-to-end human object interaction detection with hoi transformer
Cheng Zou, Bohan Wang, Yue Hu, Junqi Liu, Qian Wu, Yu Zhao, Boxun Li, Chenguang Zhang, Chi Zhang, Yichen Wei, et al · 2021
Later among the works it cites.
Bongard-hoi: Benchmarking few-shot visual reasoning for human-object interactions
Huaizu Jiang, Xiaojian Ma, Weili Nie, Zhiding Yu, Yuke Zhu, and Anima Anandkumar · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Bolya, Sean Foley, James Hays, and Judy Hoffman · 2020
Cited alongside, same era.
Smooth-ap: Smoothing the path towards large-scale image retrieval
Andrew Brown, Weidi Xie, Vicky Kalogeiton, and Andrew Zisserman · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
DRG: Dual relation graph for human-object interaction detection
Chen Gao, Jiarui Xu, Yuliang Zou, and Jia-Bin Huang · 2020
Cited alongside, same era.
Diagnosing rarity in human-object interaction detection
Mert Kilickaya and Arnold Smeulders · 2020
Cited alongside, same era.
Solar: second-order loss and attention for image retrieval
Tony Ng, Vassileios Balntas, Yurun Tian, and Krystian Mikolajczyk · 2020
Cited alongside, same era.
VSGNet: Spatial attention network for detecting human object interactions using graph convolutions
Oytun Ulutan, ASM Iftekhar, and Bangalore S Manjunath · 2020
Cited alongside, same era.
Yong-Lu Li, Hongwei Fan, Zuoyu Qiu, Yiming Dou, Liang Xu, Hao-Shu Fang, Peiyang Guo, Haisheng Su, Dongliang Wang, Wei Wu, et al · 2022
Later among the works it cites.
Gen-vlkt: Simplify association and enhance interaction understanding for hoi detection
Yue Liao, Aixi Zhang, Miao Lu, Yongliang Wang, Xiaobo Li, and Si Liu · 2022
Later among the works it cites.
Mining cross-person cues for body-part interactiveness learning in hoi detection
Xiaoqian Wu, Yong-Lu Li, Xinpeng Liu, Junyi Zhang, Yuzhe Wu, and Cewu Lu · 2022
Later among the works it cites.
Rlip: Relational language-image pre-training for human-object interaction detection
Hangjie Yuan, Jianwen Jiang, Samuel Albanie, Tao Feng, Ziyuan Huang, Dong Ni, and Mingqian Tang · 2022
Later among the works it cites.
Towards hard-positive query mining for detr-based human-object interaction detection
Xubin Zhong, Changxing Ding, Zijian Li, and Shaoli Huang · 2022
Later among the works it cites.
Relational context learning for human-object interaction detection
Sanghyun Kim, Deunsol Jung, and Minsu Cho · 2023
Closest in time.
Fgahoi: Fine-grained anchors for human-object interaction detection
Shuailei Ma, Yuefeng Wang, Shanze Wang, and Ying Wei · 2023
Closest in time.
Fine-grained affordance annotation for egocentric hand-object interaction videos
Zecheng Yu, Yifei Huang, Ryosuke Furuta, Takuma Yagi, Yusuke Goutsu, and Yoichi Sato · 2023
Closest in time.
Rlipv2: Fast scaling of relational language-image pre-training
Hangjie Yuan, Shiwei Zhang, Xiang Wang, Samuel Albanie, Yining Pan, Tao Feng, Jianwen Jiang, Dong Ni, Yingya Zhang, and Deli Zhao · 2023
Closest in time.