Fetching the paper…
Reading the bibliography…
We present a method for inferring diverse 3D models of human-object interactions from images.
“cloze procedure”: A new tool for measuring readability
Wilson L Taylor · 1953
Earlier work this paper cites.
Object properties and knowledge in early lexical learning
Susan S Jones, Linda B Smith, and Barbara Landau · 1991
Earlier work this paper cites.
Pedestrian detection from a moving vehicle
Dariu M Gavrila · 2000
Earlier work this paper cites.
Turning the tables: Language and spatial reasoning
Peggy Li and Lila Gleitman · 2002
Earlier work this paper cites.
k-means++: The advantages of careful seeding
David Arthur and Sergei Vassilvitskii · 2006
Earlier work this paper cites.
SMPL: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black · 2015
Earlier work this paper cites.
Pigraphs: learning interaction snapshots from observations
Manolis Savva, Angel X Chang, Pat Hanrahan, Matthew Fisher, and Matthias Nießner · 2016
Earlier work this paper cites.
Embodied hands: Modeling and capturing hands and bodies together
Javier Romero, Dimitrios Tzionas, and Michael J. Black · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi · 2017
Earlier work this paper cites.
Neural 3d mesh renderer
Hiroharu Kato, Yoshitaka Ushiku, and Tatsuya Harada · 2018
Earlier work this paper cites.
Matryoshka networks: Predicting 3d geometry via nested shape layers
Stephan R Richter and Stefan Roth · 2018
Earlier work this paper cites.
Resolving 3d human pose ambiguities with 3d scene constraints
Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black · 2019
Earlier work this paper cites.
Escaping plato’s cave: 3d shape from adversarial rendering
Philipp Henzler, Niloy J Mitra, and Tobias Ritschel · 2019
Earlier work this paper cites.
Natural language guided visual relationship detection
Wentong Liao, Bodo Rosenhahn, Ling Shuai, and Michael Ying Yang · 2019
Earlier work this paper cites.
Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding
Kaichun Mo, Shilin Zhu, Angel X Chang, Li Yi, Subarna Tripathi, Leonidas J Guibas, and Hao Su · 2019
Cited alongside, same era.
Expressive body capture: 3D hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Cited alongside, same era.
Neural state machine for character-scene interactions
Sebastian Starke, He Zhang, Taku Komura, and Jun Saito · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Compositional networks enable systematic generalization for grounded language understanding
Yen-Ling Kuo, Boris Katz, and Andrei Barbu · 2021
Later among the works it cites.
Clip-it! language-guided video summarization
Medhini Narasimhan, Anna Rohrbach, and Trevor Darrell · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Holistic 3d human and scene mesh estimation from single view images
Zhenzhen Weng and Serena Yeung · 2021
Later among the works it cites.
Behave: Dataset and method for tracking human object interactions
Bharat Lal Bhatnagar, Xianghui Xie, Ilya A Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MMPose Contributors · 2020
Cited alongside, same era.
Pointrend: Image segmentation as rendering
Alexander Kirillov, Yuxin Wu, Kaiming He, and Ross Girshick · 2020
Cited alongside, same era.
Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger · 2020
Cited alongside, same era.
Language-guided semantic mapping and mobile manipulation in partially observable environments
Siddharth Patki, Ethan Fahnestock, Thomas M Howard, and Matthew R Walter · 2020
Cited alongside, same era.
Perceiving 3d human-object spatial arrangements from a single image in the wild
Jason Y Zhang, Sam Pepose, Hanbyul Joo, Deva Ramanan, Jitendra Malik, and Angjoo Kanazawa · 2020
Cited alongside, same era.
Robust visual reasoning via language guided neural module networks
Arjun Akula, Varun Jampani, Soravit Changpinyo, and Song-Chun Zhu · 2021
Cited alongside, same era.
From points to multi-object 3d reconstruction
Francis Engelmann, Konstantinos Rematas, Bastian Leibe, and Vittorio Ferrari · 2021
Cited alongside, same era.
Arthur Bucker, Luis Figueredo, Sami Haddadin, Ashish Kapoor, Shuang Ma, and Rogerio Bonatti · 2022
Closest in time.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch · 2022
Closest in time.
Denseclip: Language-guided dense prediction with context-aware prompting
Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu · 2022
Closest in time.
Skill induction and planning with latent language
Pratyusha Sharma, Antonio Torralba, and Jacob Andreas · 2022
Closest in time.
Chore: Contact, human and object reconstruction from a single rgb image
Xianghui Xie, Bharat Lal Bhatnagar, and Gerard Pons-Moll · 2022
Closest in time.
Human-aware object placement for visual environment reconstruction
Hongwei Yi, Chun-Hao P Huang, Dimitrios Tzionas, Muhammed Kocabas, Mohamed Hassan, Siyu Tang, Justus Thies, and Michael J Black · 2022
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Adrian Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, et al · 2022
Closest in time.
Couch: Towards controllable human-chair interactions
Xiaohan Zhang, Bharat Lal Bhatnagar, Vladimir Guzov, Sebastian Starke, and Gerard Pons-Moll · 2022
Closest in time.