Fetching the paper…
Reading the bibliography…
Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI.
Color indexing
Michael J Swain and Dana H Ballard · 1991
Earlier work this paper cites.
What is body image?
Peter David Slade · 1994
Earlier work this paper cites.
The folk concept of intentionality
Bertram F Malle and Joshua Knobe · 1997
Earlier work this paper cites.
The coordination of eye, head, and hand movements in a natural task
Jeff Pelz, Mary Hayhoe, and Russ Loeber · 2001
Earlier work this paper cites.
Cognitive neuroscience of human social behaviour
Ralph Adolphs · 2003
Earlier work this paper cites.
Embodiment and cognitive science
Raymond W Gibbs Jr · 2005
Earlier work this paper cites.
Auc: a misleading measure of the performance of predictive distribution models
Jorge M Lobo, Alberto Jiménez-Valverde, and Raimundo Real · 2008
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
Bangpeng Yao and Li Fei-Fei · 2010
Earlier work this paper cites.
Affordances of augmented reality in science learning: Suggestions for future research
Kun-Hung Cheng and Chin-Chung Tsai · 2013
Earlier work this paper cites.
The image and appearance of the human body
Paul Schilder · 2013
Earlier work this paper cites.
Modeling 4d human-object interactions for event and object recognition
Ping Wei, Yibiao Zhao, Nanning Zheng, and Song-Chun Zhu · 2013
Earlier work this paper cites.
The ecological approach to visual perception: classic edition
James J Gibson · 2014
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi · 2016
Earlier work this paper cites.
Optimizing intersection-over-union in deep neural networks for image segmentation
Md Atiqur Rahman and Yang Wang · 2016
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Learning to detect human-object interactions
Yu-Wei Chao, Yunfan Liu, Xieyang Liu, Huayi Zeng, and Jia Deng · 2018
Earlier work this paper cites.
Demo2vec: Reasoning object affordances from online videos
Kuan Fang, Te-Lin Wu, Daniel Yang, Silvio Savarese, and Joseph J Lim · 2018
Earlier work this paper cites.
Detecting and recognizing human-object interactions
Georgia Gkioxari, Ross Girshick, Piotr Dollár, and Kaiming He · 2018
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Learning 3d human dynamics from video
Angjoo Kanazawa, Jason Y Zhang, Panna Felsen, and Jitendra Malik · 2019
Earlier work this paper cites.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Earlier work this paper cites.
Putting humans in a scene: Learning affordance in 3d indoor environments
Xueting Li, Sifei Liu, Kihwan Kim, Xiaolong Wang, Ming-Hsuan Yang, and Jan Kautz · 2019
Earlier work this paper cites.
Grounded human-object interaction hotspots from video
Tushar Nagarajan, Christoph Feichtenhofer, and Kristen Grauman · 2019
Earlier work this paper cites.
Habitat: A platform for embodied ai research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, et al · 2019
Earlier work this paper cites.
Lsta: Long short-term attention for egocentric action recognition
Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz · 2019
Earlier work this paper cites.
Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data
Mikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Thanh Nguyen, and Sai-Kit Yeung · 2019
Earlier work this paper cites.
Dynamic graph cnn for learning on point clouds
Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon · 2019
Earlier work this paper cites.
Vibe: Video inference for human body pose and shape estimation
Muhammed Kocabas, Nikos Athanasiou, and Michael J Black · 2020
Earlier work this paper cites.
Ego2hands: A dataset for egocentric two-hand segmentation and detection
Fanqing Lin, Brian Price, and Tony Martinez · 2020
Earlier work this paper cites.
Ego-topo: Environment affordances from egocentric video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman · 2020
Earlier work this paper cites.
Detecting hands and recognizing physical contact in the wild
Supreeth Narasimhaswamy, Trung Nguyen, and Minh Hoai Nguyen · 2020
Earlier work this paper cites.
Deep high-resolution representation learning for visual recognition
Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al · 2020
Earlier work this paper cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Earlier work this paper cites.
Learning to anticipate egocentric actions by imagination
Yu Wu, Linchao Zhu, Xiaohan Wang, Yi Yang, and Fei Wu · 2020
Earlier work this paper cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Earlier work this paper cites.
The epic-kitchens dataset: Collection, challenges and baselines
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2021
Earlier work this paper cites.
3d affordancenet: A benchmark for visual object affordance understanding
Shengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen, and Kui Jia · 2021
Cited alongside, same era.
Populating 3d scenes by learning human-scene interaction
Mohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, and Michael J Black · 2021
Cited alongside, same era.
Intentonomy: a dataset and study towards human intent understanding
Menglin Jia, Zuxuan Wu, Austin Reiter, Claire Cardie, Serge Belongie, and Ser-Nam Lim · 2021
Cited alongside, same era.
Learning dexterous grasping with object-centric visual affordances
Priyanka Mandikal and Kristen Grauman · 2021
Cited alongside, same era.
Where2act: From pixels to actions for articulated 3d objects
Kaichun Mo, Leonidas J Guibas, Mustafa Mukadam, Abhinav Gupta, and Shubham Tulsiani · 2021
Cited alongside, same era.
Shaping embodied agent behavior with activity-context priors from egocentric video
Object motion guided human motion synthesis
Jiaman Li, Jiajun Wu, and C Karen Liu · 2023
Later among the works it cites.
Grounded affordance from exocentric view
Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao, and Dacheng Tao · 2023
Later among the works it cites.
Intention-conditioned long-term human egocentric action anticipation
Esteve Valls Mascaró, Hyemin Ahn, and Dongheui Lee · 2023
Later among the works it cites.
Multi-label affordance mapping from egocentric vision
Lorenzo Mur-Labadia, Jose J Guerrero, and Ruben Martinez-Cantin · 2023
Later among the works it cites.
Action sensitivity learning for the ego4d episodic memory challenge 2023
Jiayi Shao, Xiaohan Wang, Ruijie Quan, and Yi Yang · 2023
Later among the works it cites.
Egodistill: Egocentric head motion distillation for efficient video understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tushar Nagarajan and Kristen Grauman · 2021
Cited alongside, same era.
Interactive prototype learning for egocentric action recognition
Xiaohan Wang, Linchao Zhu, Heng Wang, and Yi Yang · 2021
Cited alongside, same era.
Cpf: Learning a contact potential field to model the hand-object interaction
Lixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu, Jiefeng Li, and Cewu Lu · 2021
Cited alongside, same era.
Visual acoustic matching
Changan Chen, Ruohan Gao, Paul Calamia, and Kristen Grauman · 2022
Cited alongside, same era.
Soundspaces 2.0: A simulation platform for visual-acoustic learning
Changan Chen, Carl Schissler, Sanchit Garg, Philip Kobernik, Alexander Clegg, Paul Calamia, Dhruv Batra, Philip Robinson, and Kristen Grauman · 2022
Cited alongside, same era.
Masked-attention mask transformer for universal image segmentation
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar · 2022
Cited alongside, same era.
A survey of embodied ai: From simulators to research tasks
Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan · 2022
Cited alongside, same era.
Shuhan Tan, Tushar Nagarajan, and Kristen Grauman · 2023
Later among the works it cites.
Deco: Dense estimation of 3d human-scene contact in the wild
Shashank Tripathi, Agniv Chatterjee, Jean-Claude Passy, Hongwei Yi, Dimitrios Tzionas, and Michael J Black · 2023
Later among the works it cites.
Weikang Wan, Haoran Geng, Yun Liu, Zikang Shan, Yaodong Yang, Li Yi, and He Wang · 2023
Later among the works it cites.
Find what you want: Learning demand-conditioned object attribute space for demand-driven navigation
Hongcheng Wang, Andy Guan Hong Chen, Xiaoqi Li, Mingdong Wu, and Hao Dong · 2023
Later among the works it cites.
Egocentric whole-body motion capture with fisheyevit and diffusion-based motion refinement
Jian Wang, Zhe Cao, Diogo Luvizon, Lingjie Liu, Kripasindhu Sarkar, Danhang Tang, Thabo Beeler, and Christian Theobalt · 2023
Later among the works it cites.
Embodiedscan: A holistic multi-modal 3d perception suite towards embodied ai
Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu, Ruiyuan Lyu, Peisen Li, Xiao Chen, Wenwei Zhang, Kai Chen, Tianfan Xue, et al · 2023
Later among the works it cites.
Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation
Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Jiawei Ren, Liang Pan, Wayne Wu, Lei Yang, Jiaqi Wang, Chen Qian, et al · 2023
Later among the works it cites.
Interdiff: Generating 3d human-object interactions with physics-informed diffusion
Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui · 2023
Later among the works it cites.
Learning object state changes in videos: An open-world perspective
Zihui Xue, Kumar Ashutosh, and Kristen Grauman · 2023
Later among the works it cites.
Egocentric video task translation
Zihui Xue, Yale Song, Kristen Grauman, and Lorenzo Torresani · 2023
Later among the works it cites.
Grounding 3d object affordance from 2d interactions in images
Yuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao, Jiebo Luo, and Zheng-Jun Zha · 2023
Later among the works it cites.
Fine-grained affordance annotation for egocentric hand-object interaction videos
Zecheng Yu, Yifei Huang, Ryosuke Furuta, Takuma Yagi, Yusuke Goutsu, and Yoichi Sato · 2023
Later among the works it cites.
Probabilistic human mesh recovery in 3d scenes from egocentric views
Siwei Zhang, Qianli Ma, Yan Zhang, Sadegh Aliakbarian, Darren Cosker, and Siyu Tang · 2023
Later among the works it cites.
Synthesizing diverse human motions in 3d indoor scenes
Kaifeng Zhao, Yan Zhang, Shaofei Wang, Thabo Beeler, and Siyu Tang · 2023
Later among the works it cites.
Learning video representations from large language models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar · 2023
Later among the works it cites.
Egoobjects: A large-scale egocentric dataset for fine-grained object understanding
Chenchen Zhu, Fanyi Xiao, Andrés Alvarado, Yasmine Babaei, Jiabo Hu, Hichem El-Mohri, Sean Culatana, Roshan Sumbaly, and Zhicheng Yan · 2023
Later among the works it cites.
Diff-lfd: Contact-aware model-based learning from visual demonstration for robotic manipulation via differentiable physics-based simulation and rendering
Xinghao Zhu, JingHan Ke, Zhixuan Xu, Zhixin Sun, Bizhe Bai, Jun Lv, Qingtao Liu, Yuwei Zeng, Qi Ye, Cewu Lu, et al · 2023
Later among the works it cites.
Smpler-x: Scaling up expressive human pose and shape estimation
Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei, Qingping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Mingyuan Zhang, Lei Zhang, et al · 2024
Closest in time.
Soundingactions: Learning how actions sound from narrated egocentric videos
Changan Chen, Kumar Ashutosh, Rohit Girdhar, David Harwath, and Kristen Grauman · 2024
Closest in time.
Simpleego: Predicting probabilistic body pose from egocentric cameras
Hanz Cuevas-Velasquez, Charlie Hewitt, Sadegh Aliakbarian, and Tadas Baltrušaitis · 2024
Closest in time.
Scenefun3d: Fine-grained functionality and affordance understanding in 3d scenes
Alexandros Delitzas, Ayca Takmaz, Federico Tombari, Robert Sumner, Marc Pollefeys, and Francis Engelmann · 2024
Closest in time.
Scaling up dynamic human-scene interaction modeling
Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang · 2024
Closest in time.
Zero-shot learning for the primitives of 3d affordance in general objects
Hyeonwoo Kim, Sookwan Han, Patrick Kwon, and Hanbyul Joo · 2024
Closest in time.
EgoGen: An Egocentric Synthetic Data Generator
Gen Li, Kaifeng Zhao, Siwei Zhang, Xiaozhong Lyu, Mihai Dusmanu, Yan Zhang, Marc Pollefeys, and Siyu Tang · 2024
Closest in time.
Egoenv: Human-centric environment representations from egocentric video
Tushar Nagarajan, Santhosh Kumar Ramakrishnan, Ruta Desai, James Hillis, and Kristen Grauman · 2024
Closest in time.
Joint reconstruction of 3d human and object via contact-based refinement transformer
Hyeongjin Nam, Daniel Sungho Jung, Gyeongsik Moon, and Kyoung Mu Lee · 2024
Closest in time.
A backpack full of skills: Egocentric video understanding with diverse task perspectives, 2024
Simone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, and Giuseppe Averta · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al · 2024
Closest in time.
Interaction region visual transformer for egocentric action anticipation
Debaditya Roy, Ramanathan Rajendiran, and Basura Fernando · 2024
Closest in time.
Self-supervised visual acoustic matching
Arjun Somayazulu, Changan Chen, and Kristen Grauman · 2024
Closest in time.
Lemon: Learning 3d human-object interaction relation from 2d images
Yuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao, and Zheng-Jun Zha · 2024
Closest in time.
Unimd: Towards unifying moment retrieval and temporal action detection
Yingsen Zeng, Yujie Zhong, Chengjian Feng, and Lin Ma · 2024
Closest in time.
Bidirectional progressive transformer for interaction intention anticipation
Zichen Zhang, Hongchen Luo, Wei Zhai, Yang Cao, and Yu Kang · 2024
Closest in time.