Fetching the paper…
Reading the bibliography…
Naturally controllable human-scene interaction (HSI) generation has an important role in various fields, such as VR/AR content creation and human-centered AI.
Shape2pose: Human-centric shape analysis
Vladimir G Kim, Siddhartha Chaudhuri, Leonidas Guibas, and Thomas Funkhouser · 2014
Earlier work this paper cites.
Indexing 3D scenes using the interaction bisector surface
Xi Zhao, He Wang, and Taku Komura · 2014
Earlier work this paper cites.
Pigraphs: learning interaction snapshots from observations
Manolis Savva, Angel X Chang, Pat Hanrahan, Matthew Fisher, and Matthias Nießner · 2016
Earlier work this paper cites.
Matterport3D: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3D reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Text2action: Generative adversarial synthesis from language to action
Hyemin Ahn, Timothy Ha, Yunho Choi, Hwiyeon Yoo, and Songhwai Oh · 2018
Earlier work this paper cites.
Where and who? automatic semantic-aware person composition
Fuwen Tan, Crispin Bernier, Benjamin Cohen, Vicente Ordonez, and Connelly Barnes · 2018
Earlier work this paper cites.
Language2pose: Natural language grounded pose forecasting
Chaitanya Ahuja and Louis-Philippe Morency · 2019
Earlier work this paper cites.
Holistic++ scene understanding: Single-view 3D holistic scene parsing and human pose estimation with human-object interaction and physical commonsense
Yixin Chen, Siyuan Huang, Tao Yuan, Siyuan Qi, Yixin Zhu, and Song-Chun Zhu · 2019
Earlier work this paper cites.
Resolving 3D human pose ambiguities with 3D scene constraints
Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black · 2019
Earlier work this paper cites.
Deepgcns: Can gcns go as deep as cnns?
Guohao Li, Matthias Muller, Ali Thabet, and Bernard Ghanem · 2019
Earlier work this paper cites.
Putting humans in a scene: Learning affordance in 3D indoor environments
Xueting Li, Sifei Liu, Kihwan Kim, Xiaolong Wang, Ming-Hsuan Yang, and Jan Kautz · 2019
Earlier work this paper cites.
imapper: interaction-guided scene mapping from monocular videos
Aron Monszpart, Paul Guerrero, Duygu Ceylan, Ersin Yumer, and Niloy J Mitra · 2019
Earlier work this paper cites.
Expressive body capture: 3D hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black · 2019
Cited alongside, same era.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma · 2019
Cited alongside, same era.
On the continuity of rotation representations in neural networks
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li · 2019
Cited alongside, same era.
Scenegraphnet: Neural message passing for 3D indoor scene augmentation
Yang Zhou, Zachary While, and Evangelos Kalogerakis · 2019
Cited alongside, same era.
Referit3D: Neural listeners for fine-grained 3D object identification in real-world scenes
Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas Guibas · 2020
Cited alongside, same era.
Synthesizing long-term 3D human motion and interaction in 3D scenes
Jiashun Wang, Huazhe Xu, Jingwei Xu, Sifei Liu, and Xiaolong Wang · 2021
Later among the works it cites.
Sat: 2D semantics assisted training for 3D visual grounding
Zhengyuan Yang, Songyang Zhang, Liwei Wang, and Jiebo Luo · 2021
Later among the works it cites.
Teach: Temporal action composition for 3D humans
Nikos Athanasiou, Mathis Petrovich, Michael J Black, and Gül Varol · 2022
Later among the works it cites.
Generating diverse and natural 3D human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng · 2022
Later among the works it cites.
Avatarclip: Zero-shot text-driven generation and animation of 3D avatars
Fangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai, Lei Yang, and Ziwei Liu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scanrefer: 3D object localization in rgb-d scans using natural language
Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner · 2020
Cited alongside, same era.
Learning 3D semantic scene graphs from 3D indoor reconstructions
Johanna Wald, Helisa Dhamo, Nassir Navab, and Federico Tombari · 2020
Cited alongside, same era.
Place: Proximity learning of articulation and contact in 3D environments
Siwei Zhang, Yan Zhang, Qianli Ma, Michael J Black, and Siyu Tang · 2020
Cited alongside, same era.
Generating 3D people in scenes without people
Yan Zhang, Mohamed Hassan, Heiko Neumann, Michael J Black, and Siyu Tang · 2020
Cited alongside, same era.
Free-form description guided 3D visual graph network for object grounding in point cloud
Mingtao Feng, Zhen Li, Qi Li, Liang Zhang, XiangDong Zhang, Guangming Zhu, Hui Zhang, Yaonan Wang, and Ajmal Mian · 2021
Cited alongside, same era.
Synthesis of compositional animations from textual descriptions
Anindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt, and Philipp Slusallek · 2021
Cited alongside, same era.
Populating 3D scenes by learning human-scene interaction
Mohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, and Michael J Black · 2021
Cited alongside, same era.
3D-sps: Single-stage 3D visual grounding via referred point progressive selection
Junyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao, Haibing Ren, Hao Shen, Huaxia Xia, and Si Liu · 2022
Later among the works it cites.
Temos: Generating diverse human motions from textual descriptions
Mathis Petrovich, Michael J Black, and Gül Varol · 2022
Later among the works it cites.
Languagerefer: Spatial-language model for 3D visual grounding
Junha Roh, Karthik Desingh, Ali Farhadi, and Dieter Fox · 2022
Later among the works it cites.
Actformer: A gan transformer framework towards general action-conditioned 3D human motion generation
Ziyang Song, Dongliang Wang, Nan Jiang, Zhicheng Fang, Chenjing Ding, Weihao Gan, and Wei Wu · 2022
Later among the works it cites.
Towards diverse and natural scene-aware 3D human motion synthesis
Jingbo Wang, Yu Rong, Jingyuan Liu, Sijie Yan, Dahua Lin, and Bo Dai · 2022
Later among the works it cites.
Clip-actor: Text-driven recommendation and stylization for animating human meshes
Kim Youwang, Kim Ji-Yeon, and Tae-Hyun Oh · 2022
Later among the works it cites.
Compositional human-scene interaction synthesis with semantic control
Kaifeng Zhao, Shaofei Wang, Yan Zhang, Thabo Beeler, and Siyu Tang · 2022
Later among the works it cites.
Diffusion-based generation, optimization, and planning in 3D scenes
Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu · 2023
Closest in time.