Fetching the paper…
Reading the bibliography…
We present DreamHOI, a novel method for zero-shot synthesis of human-object interactions (HOIs), enabling a 3D human model to realistically interact with any given object based on a textual description.
Abstract muscle action procedures for human face animation
Nadia Magnenat-Thalmann, E Primeau, and Daniel Thalmann · 1988
Earlier work this paper cites.
Layered construction for deformable animated characters
John E Chadwick, David R Haumann, and Richard E Parent · 1989
Earlier work this paper cites.
Embedded deformation for shape manipulation
Robert W Sumner, Johannes Schmid, and Mark Pauly · 2007
Earlier work this paper cites.
SMPL: a skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black · 2015
Earlier work this paper cites.
Convolutional pose machines
Shih-En Wei, Varun Ramakrishna, Takeo Kanade, and Yaser Sheikh · 2016
Earlier work this paper cites.
Realtime multi-person 2d pose estimation using part affinity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh · 2017
Earlier work this paper cites.
Embodied hands: Modeling and capturing hands and bodies together
Javier Romero, Dimitrios Tzionas, and Michael J. Black · 2017
Earlier work this paper cites.
Hand keypoint detection in single images using multiview bootstrapping
Tomas Simon, Hanbyul Joo, Iain Matthews, and Yaser Sheikh · 2017
Earlier work this paper cites.
3D menagerie: Modeling the 3D shape and pose of animals
Silvia Zuffi, Angjoo Kanazawa, David W. Jacobs, and Michael J. Black · 2017
Earlier work this paper cites.
Pixel2mesh: Generating 3d mesh models from single rgb images
Nanyang Wang, Yinda Zhang, Zhuwen Li, Yanwei Fu, Wei Liu, and Yu-Gang Jiang · 2018
Earlier work this paper cites.
Expressive body capture: 3D hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black · 2019
Earlier work this paper cites.
Deepvoxels: Learning persistent 3d feature embeddings
Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, and Michael Zollhofer · 2019
Earlier work this paper cites.
Compositional visual generation with energy based models
Yilun Du, Shuang Li, and Igor Mordatch · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Self-supervised learning of interpretable keypoints from unlabelled videos
Tomas Jakab, Ankush Gupta, Hakan Bilen, and Andrea Vedaldi · 2020
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Earlier work this paper cites.
STAR: A sparse trained articulated human body regressor
Ahmed A A Osman, Timo Bolkart, and Michael J. Black · 2020
Earlier work this paper cites.
Grab: A dataset of whole-body human grasping of objects
Omid Taheri, Nima Ghorbani, Michael J Black, and Dimitrios Tzionas · 2020
Earlier work this paper cites.
Unsupervised learning of probably symmetric deformable 3d objects from images in the wild
Shangzhe Wu, Christian Rupprecht, and Andrea Vedaldi · 2020
Earlier work this paper cites.
Generating 3d people in scenes without people
Yan Zhang, Mohamed Hassan, Heiko Neumann, Michael J Black, and Siyu Tang · 2020
Earlier work this paper cites.
Populating 3d scenes by learning human-scene interaction
Mohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, and Michael J Black · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
D3d-hoi: Dynamic 3d human-object interactions from videos
Xiang Xu, Hanbyul Joo, Greg Mori, and Manolis Savva · 2021
Cited alongside, same era.
pixelnerf: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa · 2021
Cited alongside, same era.
Pamir: Parametric model-conditioned implicit representation for image-based human reconstruction
Zerong Zheng, Tao Yu, Yebin Liu, and Qionghai Dai · 2021
Cited alongside, same era.
Behave: Dataset and method for tracking human object interactions
Bharat Lal Bhatnagar, Xianghui Xie, Ilya Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll · 2022
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning, 2022
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2022
Cited alongside, same era.
Cg3d: Compositional generation for text-to-3d via gaussian splatting
Alexander Vilesov, Pradyumna Chari, and Achuta Kadambi · 2023
Later among the works it cites.
Decoupling human and camera motion from videos in the wild
Vickie Ye, Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa · 2023
Later among the works it cites.
Scenewiz3d: Towards text-guided 3d scene composition
Qihang Zhang, Chaoyang Wang, Aliaksandr Siarohin, Peiye Zhuang, Yinghao Xu, Ceyuan Yang, Dahua Lin, Bolei Zhou, Sergey Tulyakov, and Hsin-Ying Lee · 2023
Later among the works it cites.
Comboverse: Compositional 3d assets creation using spatially-aware diffusion guidance
Yongwei Chen, Tengfei Wang, Tong Wu, Xingang Pan, Kui Jia, and Ziwei Liu · 2024
Closest in time.
Interfusion: Text-driven generation of 3d human-object interaction
Sisi Dai, Wenhao Li, Haowen Sun, Haibin Huang, Chongyang Ma, Hui Huang, Kai Xu, and Ruizhen Hu · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
Compositional visual generation with composable diffusion models
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano · 2022
Cited alongside, same era.
Couch: Towards controllable human-chair interactions
Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Vladimir Guzov, and Gerard Pons-Moll · 2022
Cited alongside, same era.
Compositional human-scene interaction synthesis with semantic control
Kaifeng Zhao, Shaofei Wang, Yan Zhang, Thabo Beeler, and Siyu Tang · 2022
Cited alongside, same era.
SMPLitex: A Generative Model and Dataset for 3D Human Texture Estimation from Single Image
Dan Casas and Marc Comino-Trinidad · 2023
Cited alongside, same era.
Closest in time.
Cg-hoi: Contact-guided 3d human-object interaction generation
Christian Diller and Angela Dai · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach · 2024
Closest in time.
Cat3d: Create anything in 3d with multi-view diffusion models
Ruiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul Srinivasan, Jonathan T Barron, and Ben Poole · 2024
Closest in time.
Farm3D: Learning articulated 3d animals by distilling 2d diffusion
Tomas Jakab, Ruining Li, Shangzhe Wu, Christian Rupprecht, and Andrea Vedaldi · 2024
Closest in time.
Im-3d: Iterative multiview diffusion and reconstruction for high-quality 3d generation
Luke Melas-Kyriazi, Iro Laina, Christian Rupprecht, Natalia Neverova, Andrea Vedaldi, Oran Gafni, and Filippos Kokkinos · 2024
Closest in time.
Compositional 3d scene generation using locally conditioned diffusion
Ryan Po and Gordon Wetzstein · 2024
Closest in time.
MVDream: Multi-view diffusion for 3D generation
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang · 2024
Closest in time.
Sketchfab
Sketchfab · 2024
Closest in time.
Splatter image: Ultra-fast single-view 3d reconstruction
Stanislaw Szymanowicz, Chrisitian Rupprecht, and Andrea Vedaldi · 2024
Closest in time.
Dreamgaussian: Generative gaussian splatting for efficient 3d content creation
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng · 2024
Closest in time.
Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion
Vikram Voleti, Chun-Han Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, and Varun Jampani · 2024
Closest in time.
Tram: Global trajectory and motion of 3d humans from in-the-wild videos
Yufu Wang, Ziyun Wang, Lingjie Liu, and Kostas Daniilidis · 2024
Closest in time.
Thor: Text to human-object interaction diffusion via relation intervention
Qianyang Wu, Ye Shi, Xiaoshui Huang, Jingyi Yu, Lan Xu, and Jingya Wang · 2024
Closest in time.
Free3d: Consistent novel view synthesis without 3d representation
Chuanxia Zheng and Andrea Vedaldi · 2024
Closest in time.
Gala3d: Towards text-to-3d complex scene generation via layout-guided generative gaussian splatting
Xiaoyu Zhou, Xingjian Ran, Yajiao Xiong, Jinlin He, Zhiwei Lin, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang · 2024
Closest in time.
Varen: Very accurate and realistic equine network
Silvia Zuffi, Ylva Mellbin, Ci Li, Markus Hoeschle, Hedvig Kjellström, Senya Polikovsky, Elin Hernlund, and Michael J Black · 2024
Closest in time.