Fetching the paper…
Reading the bibliography…
Generating natural hand-object interactions in 3D is challenging as the resulting hand and object motions are expected to be physically plausible and semantically meaningful.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In Computer Vision and Pattern Recognition (CVPR) . Ieee, 248–255
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Synthesis of detailed hand manipulations using contact sampling
Yuting Ye and C Karen Liu. 2012 · 2012
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning . PMLR, 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Embodied Hands: Modeling and Capturing Hands and Bodies Together
Javier Romero, Dimitrios Tzionas, and Michael J. Black. 2017 · 2017
Earlier work this paper cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations. In Robotics: Science and Systems (RSS)
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine. 2018 · 2018
Earlier work this paper cites.
Contactdb: Analyzing and predicting grasp contact via thermal imaging. In Computer Vision and Pattern Recognition (CVPR) . 8709–8719
Samarth Brahmbhatt, Cusuh Ham, Charles C Kemp, and James Hays. 2019 · 2019
Earlier work this paper cites.
Learning joint reconstruction of hands and manipulated objects. In Computer Vision and Pattern Recognition (CVPR)
Yana Hasson, Gül Varol, Dimitris Tzionas, Igor Kalevatykh, Michael J. Black, Ivan Laptev, and Cordelia Schmid. 2019 · 2019
Earlier work this paper cites.
AMASS: Archive of Motion Capture as Surface Shapes. In International Conference on Computer Vision . 5442–5451
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. 2019 · 2019
Earlier work this paper cites.
Image generation from small datasets via batch statistics adaptation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 2750–2758
Atsuhiro Noguchi and Tatsuya Harada. 2019 · 2019
Earlier work this paper cites.
Efficient learning on point clouds with basis point sets. In Computer Vision and Pattern Recognition (CVPR) . 4332–4341
Sergey Prokudin, Christoph Lassner, and Javier Romero. 2019 · 2019
Earlier work this paper cites.
ContactPose: A dataset of grasps with object contact and hand pose. In European Conference on Computer Vision (ECCV) . Springer, 361–378
Samarth Brahmbhatt, Chengcheng Tang, Christopher D Twigg, Charles C Kemp, and James Hays. 2020 · 2020
Earlier work this paper cites.
Effectively unbiased fid and inception score and where to find them. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 6070–6079
Min Jin Chong and David Forsyth. 2020 · 2020
Earlier work this paper cites.
Physics-based dexterous manipulations with estimated hand poses and residual reinforcement learning. In International Conference on Robotics and Automation (ICRA) . IEEE, 9561–9568
Guillermo Garcia-Hernando, Edward Johns, and Tae-Kyun Kim. 2020 · 2020
Earlier work this paper cites.
HOnnotate: A method for 3D Annotation of Hand and Object Poses. In Computer Vision and Pattern Recognition (CVPR)
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Grasping field: Learning implicit representations for human grasps. In 2020 International Conference on 3D Vision (3DV) . IEEE, 333–344
Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. 2020 · 2020
Earlier work this paper cites.
GRAB: A Dataset of Whole-Body Human Grasping of Objects. In European Conference on Computer Vision (ECCV)
Omid Taheri, Nima Ghorbani, Michael J. Black, and Dimitrios Tzionas. 2020 · 2020
Earlier work this paper cites.
DexYCB: A Benchmark for Capturing Hand Grasping of Objects. In Computer Vision and Pattern Recognition (CVPR)
Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Dieter Fox. 2021 · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. 2021 · 2021
Earlier work this paper cites.
DiffWave: A Versatile Diffusion Model for Audio Synthesis. In International Conference on Learning Representations
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. 2021 · 2021
Earlier work this paper cites.
H2O: Two Hands Manipulating Objects for First Person Interaction Recognition. In International Conference on Computer Vision (ICCV) . 10138–10148
Taein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo, and Marc Pollefeys. 2021 · 2021
Earlier work this paper cites.
Synthesizing diverse and physically stable grasps with arbitrary hand structures using differentiable force closure estimator
Tengyu Liu, Zeyu Liu, Ziyuan Jiao, Yixin Zhu, and Song-Chun Zhu. 2021 · 2021
Earlier work this paper cites.
Dynamics-regulated kinematic policy for egocentric pose estimation
Zhengyi Luo, Ryo Hachiuma, Ye Yuan, and Kris Kitani. 2021 · 2021
Earlier work this paper cites.
Learning Dexterous Grasping with Object-Centric Visual Affordances. In International Conference on Robotics and Automation (ICRA)
Priyanka Mandikal and Kristen Grauman. 2021 · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models. In International Conference on Machine Learning . PMLR, 8162–8171
Alexander Quinn Nichol and Prafulla Dhariwal. 2021 · 2021
Earlier work this paper cites.
DexMV: Imitation Learning for Dexterous Manipulation from Human Videos
Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Scene-aware generative network for human motion synthesis. In Computer Vision and Pattern Recognition (CVPR) . 12206–12215
Jingbo Wang, Sijie Yan, Bo Dai, and Dahua Lin. 2021 · 2021
Cited alongside, same era.
Manipnet: neural manipulation synthesis with a hand-object spatial representation
He Zhang, Yuting Ye, Takaaki Shiratori, and Taku Komura. 2021 · 2021
Cited alongside, same era.
Teach: Temporal action composition for 3d humans. In International Conference on 3D Vision (3DV) . IEEE, 414–423
Locomotion-Action-Manipulation: Synthesizing Human-Scene Interactions in Complex 3D Environments
Jiye Lee and Hanbyul Joo. 2023 · 2023
Later among the works it cites.
Rosario Leonardi, Antonino Furnari, Francesco Ragusa, and Giovanni Maria Farinella. 2023 · 2023
Later among the works it cites.
Controllable Human-Object Interaction Synthesis
Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, and C Karen Liu. 2023 · 2023
Later among the works it cites.
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nikos Athanasiou, Mathis Petrovich, Michael J Black, and Gül Varol. 2022 · 2022
Cited alongside, same era.
D-Grasp: Physically Plausible Dynamic Grasp Synthesis for Hand-Object Interactions. In Computer Vision and Pattern Recognition (CVPR)
Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. 2022 · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Cited alongside, same era.
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022 · 2022
Cited alongside, same era.
aitviewer
Manuel Kaufmann, Velko Vechev, and Dario Mylonopoulos. 2022 · 2022
Cited alongside, same era.
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction. In Computer Vision and Pattern Recognition (CVPR) . 21013–21022
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Cited alongside, same era.
Learning High-DOF Reaching-and-Grasping via Dynamic Representation of Gripper-Object Interaction
Qijin She, Ruizhen Hu, Juzhan Xu, Min Liu, Kai Xu, and Hui Huang. 2022 · 2022
Cited alongside, same era.
Soshi Shimada, Franziska Mueller, Jan Bednarik, Bardia Doosti, Bernd Bickel, Danhang Tang, Vladislav Golyanik, Jonathan Taylor, Christian Theobalt, and Thabo Beeler. 2023 · 2023
Later among the works it cites.
In-Style: Bridging Text and Uncurated Videos with Style Transfer for Text-Video Retrieval
Nina Shvetsova, Anna Kukleva, Bernt Schiele, and Hilde Kuehne. 2023 · 2023
Later among the works it cites.
FLEX: Full-Body Grasping Without Full-Body Grasps. In Computer Vision and Pattern Recognition (CVPR)
Purva Tendulkar, Dídac Surís, and Carl Vondrick. 2023 · 2023
Later among the works it cites.
Human Motion Diffusion Model. In The Eleventh International Conference on Learning Representations
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. 2023 · 2023
Later among the works it cites.
Weikang Wan, Haoran Geng, Yun Liu, Zikang Shan, Yaodong Yang, Li Yi, and He Wang. 2023 · 2023
Later among the works it cites.
DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation. In International Conference on Robotics and Automation (ICRA) . IEEE, 11359–11366
Ruicheng Wang, Jialiang Zhang, Jiayi Chen, Yinzhen Xu, Puhao Li, Tengyu Liu, and He Wang. 2023 · 2023
Later among the works it cites.
Diffusion models: A comprehensive survey of methods and applications
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023 · 2023
Later among the works it cites.
PhysDiff: Physics-Guided Human Motion Diffusion Model. In International Conference on Computer Vision (ICCV)
Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. 2023 · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3836–3847
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023 · 2023
Later among the works it cites.
CAMS: CAnonicalized Manipulation Spaces for Category-Level Functional Hand-Object Manipulation Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 585–594
Juntian Zheng, Qingyuan Zheng, Lixing Fang, Yun Liu, and Li Yi. 2023 · 2023
Later among the works it cites.
On the continuity of rotation representations in neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5745–5753
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. 2019 · 2023
Later among the works it cites.
Physically Plausible Full-Body Hand-Object Interaction Synthesis. In International Conference on 3D Vision (3DV)
Jona Braun, Sammy Christen, Muhammed Kocabas, Emre Aksan, and Otmar Hilliges. 2024 · 2024
Closest in time.
HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from Video. In Computer Vision and Pattern Recognition (CVPR)
Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Muhammed Kocabas, Xu Chen, Michael J Black, and Otmar Hilliges. 2024 · 2024
Closest in time.
Optimizing Diffusion Noise Can Serve As Universal Motion Priors. In Computer Vision and Pattern Recognition (CVPR)
Korrawe Karunratanakul, Konpat Preechakul, Emre Aksan, Thabo Beeler, Supasorn Suwajanakorn, and Siyu Tang. 2024 · 2024
Closest in time.
Grasp multiple objects with one hand
Yuyang Li, Bo Liu, Yiran Geng, Puhao Li, Yaodong Yang, Yixin Zhu, Tengyu Liu, and Siyuan Huang. 2024a · 2024
Closest in time.
GeneOH Diffusion: Towards Generalizable Hand-Object Interaction Denoising via Denoising Diffusion. In The Twelfth International Conference on Learning Representations
Xueyi Liu and Li Yi. 2024 · 2024
Closest in time.
Reconstructing Hands in 3D with Transformers. In Computer Vision and Pattern Recognition (CVPR)
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. 2024 · 2024
Closest in time.
GRIP: Generating Interaction Poses Using Latent Consistency and Spatial Cues. In International Conference on 3D Vision (3DV)
Omid Taheri, Yi Zhou, Dimitrios Tzionas, Yang Zhou, Duygu Ceylan, Soren Pirk, and Michael J. Black. 2024 · 2024
Closest in time.
OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion. In Computer Vision and Pattern Recognition (CVPR) . 445–456
Xinyu Zhan, Lixin Yang, Yifei Zhao, Kangrui Mao, Hanlin Xu, Zenan Lin, Kailin Li, and Cewu Lu. 2024 · 2024
Closest in time.
GraspXL: Generating Grasping Motions for Diverse Objects at Scale
Hui Zhang, Sammy Christen, Zicong Fan, Otmar Hilliges, and Jie Song. 2024a · 2024
Closest in time.
Action2motion: Conditioned generation of 3d human motions. In Proceedings of the 28th ACM International Conference on Multimedia . 2021–2029
Chuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng. 2020 · 2029
Closest in time.