Fetching the paper…
Reading the bibliography…
This paper investigates the problem of the current HOI detection methods and introduces DiffHOI, a novel HOI detection scheme grounded on a pre-trained text-image diffusion model, which enhances the detector's performance via improved data diversity and HOI representation.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Saurabh Gupta and Jitendra Malik · 2015
Earlier work this paper cites.
Future frame prediction for anomaly detection–a new baseline
Wen Liu, Weixin Luo, Dongze Lian, and Shenghua Gao · 2018
Earlier work this paper cites.
Learning to detect human-object interactions
Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng · 2018
Earlier work this paper cites.
ican: Instance-centric attention network for human-object interaction detection
Chen Gao, Yuliang Zou, and Jia-Bin Huang · 2018
Earlier work this paper cites.
Learning human-object interactions by graph parsing neural networks
Siyuan Qi, Wenguan Wang, Baoxiong Jia, Jianbing Shen, and Song-Chun Zhu · 2018
Earlier work this paper cites.
Pose-aware multi-level feature network for human object interaction detection
Bo Wan, Desen Zhou, Yongfei Liu, Rongjie Li, and Xuming He · 2019
Earlier work this paper cites.
No-frills human-object interaction detection: Factorization, layout encodings, and training techniques
Tanmay Gupta, Alexander Schwing, and Derek Hoiem · 2019
Earlier work this paper cites.
Transferable interactiveness knowledge for human-object interaction detection
Yong-Lu Li, Siyuan Zhou, Xijie Huang, Liang Xu, Ze Ma, Hao-Shu Fang, Yanfeng Wang, and Cewu Lu · 2019
Earlier work this paper cites.
Learning to detect human-object interactions with knowledge
Bingjie Xu, Yongkang Wong, Junnan Li, Qi Zhao, and Mohan S Kankanhalli · 2019
Earlier work this paper cites.
Detecting unseen visual relations using analogies
Julia Peyre, Ivan Laptev, Cordelia Schmid, and Josef Sivic · 2019
Earlier work this paper cites.
Visual compositional learning for human-object interaction detection
Zhi Hou, Xiaojiang Peng, Yu Qiao, and Dacheng Tao · 2020
Earlier work this paper cites.
Ppdm: Parallel point detection and matching for real-time human-object interaction detection
Yue Liao, Si Liu, Fei Wang, Yanjie Chen, Chen Qian, and Jiashi Feng · 2020
Earlier work this paper cites.
Uniondet: Union-level detector towards real-time human-object interaction detection
Bumsoo Kim, Taeho Choi, Jaewoo Kang, and Hyunwoo J Kim · 2020
Earlier work this paper cites.
Detecting human-object interactions via functional generalization
Ankan Bansal, Sai Saketh Rambhatla, Abhinav Shrivastava, and Rama Chellappa · 2020
Earlier work this paper cites.
Detailed 2d-3d joint representation for human-object interaction
Yong-Lu Li, Xinpeng Liu, Han Lu, Shiyi Wang, Junqi Liu, Jiefeng Li, and Cewu Lu · 2020
Earlier work this paper cites.
Consnet: Learning consistency graph for zero-shot human-object interaction detection
Ye Liu, Junsong Yuan, and Chang Wen Chen · 2020
Cited alongside, same era.
Drg: Dual relation graph for human-object interaction detection
Chen Gao, Jiarui Xu, Yuliang Zou, and Jia-Bin Huang · 2020
Cited alongside, same era.
Contextual heterogeneous graph network for human-object interaction detection
Hai Wang, Wei-shi Zheng, and Ling Yingbiao · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
What the daam: Interpreting stable diffusion using cross attention
Raphael Tang, Akshat Pandey, Zhiying Jiang, Gefei Yang, Karun Kumar, Jimmy Lin, and Ferhan Ture · 2022
Later among the works it cites.
The overlooked classifier in human-object interaction recognition
Ying Jin, Yinpeng Chen, Lijuan Wang, Jianfeng Wang, Pei Yu, Lin Liang, Jenq-Neng Hwang, and Zicheng Liu · 2022
Later among the works it cites.
Mstr: Multi-scale transformer for end-to-end human-object interaction detection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dirv: Dense interaction region voting for end-to-end human-object interaction detection
Hao-Shu Fang, Yichen Xie, Dian Shao, and Cewu Lu · 2021
Cited alongside, same era.
Hotr: End-to-end human-object interaction detection with transformers
Bumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim, and Hyunwoo J Kim · 2021
Cited alongside, same era.
Qahoi: query-based anchors for human-object interaction detection
Junwen Chen and Keiji Yanai · 2021
Cited alongside, same era.
Reformulating hoi detection as adaptive set prediction
Mingfei Chen, Yue Liao, Si Liu, Zhiyuan Chen, Fei Wang, and Chen Qian · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Cited alongside, same era.
Bumsoo Kim, Jonghwan Mun, Kyoung-Woon On, Minchul Shin, Junhyun Lee, and Eun-Sol Kim · 2022
Later among the works it cites.
What to look at and where: Semantic and spatial refined transformer for detecting human-object interactions
ASM Iftekhar, Hao Chen, Kaustav Kundu, Xinyu Li, Joseph Tighe, and Davide Modolo · 2022
Later among the works it cites.
Distillation using oracle queries for transformer-based human-object interaction detection
Xian Qu, Changxing Ding, Xingao Li, Xubin Zhong, and Dacheng Tao · 2022
Later among the works it cites.
Interactiveness field in human-object interactions
Xinpeng Liu, Yong-Lu Li, Xiaoqian Wu, Yu-Wing Tai, Cewu Lu, and Chi-Keung Tang · 2022
Later among the works it cites.
Open-vocabulary panoptic segmentation with text-to-image diffusion models
Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiaolong Wang, and Shalini De Mello · 2023
Closest in time.
Guiding text-to-image diffusion model towards grounded generation
Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie · 2023
Closest in time.
Jeeseung Park, Jin-Woo Park, and Jong-Seok Lee · 2023
Closest in time.
Fgahoi: Fine-grained anchors for human-object interaction detection
Shuailei Ma, Yuefeng Wang, Shanze Wang, and Ying Wei · 2023
Closest in time.
Hoiclip: Efficient knowledge transfer for hoi detection with vision-language models
Shan Ning, Longtian Qiu, Yongfei Liu, and Xuming He · 2023
Closest in time.
Effective data augmentation with diffusion models
Brandon Trabucco, Kyle Doherty, Max Gurinas, and Ruslan Salakhutdinov · 2023
Closest in time.
Leaving reality to imagination: Robust classification via generated datasets
Hritik Bansal and Aditya Grover · 2023
Closest in time.
Diffusion models and semi-supervised learners benefit mutually with few labels
Zebin You, Yong Zhong, Fan Bao, Jiacheng Sun, Chongxuan Li, and Jun Zhu · 2023
Closest in time.