Fetching the paper…
Reading the bibliography…
Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characteristics of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality and lack semantics.
Geometric pose affordance: 3d human pose with scene constraints
Wang, Z., Chen, L., Rathore, S., Shin, D., and Fowlkes, C. (2019) · 1905
Earlier work this paper cites.
The replica dataset: A digital replica of indoor spaces
Straub, J., Whelan, T., Ma, L., Chen, Y., Wijmans, E., Green, S., Engel, J. J., Mur-Artal, R., Ren, C., Verma, S., et al. (2019) · 1906
Earlier work this paper cites.
Learning to sit: Synthesizing human-chair interactions via hierarchical control
Chao, Y.-W., Yang, J., Chen, W., and Deng, J. (2019) · 1908
Earlier work this paper cites.
Multidimensional binary search trees used for associative searching
Bentley, J. L. (1975) · 1975
Earlier work this paper cites.
Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Ionescu, C., Papava, D., Olaru, V., and Sminchisescu, C. (2013) · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
Sohn, K., Lee, H., and Yan, X. (2015) · 2015
Earlier work this paper cites.
Pigraphs: learning interaction snapshots from observations
Savva, M., Chang, A. X., Hanrahan, P., Fisher, M., and Nießner, M. (2016) · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Dai, A., Chang, A. X., Savva, M., Halber, M., Funkhouser, T., and Nießner, M. (2017) · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Qi, C. R., Yi, L., Su, H., and Guibas, L. J. (2017) · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., and Van Den Hengel, A. (2018) · 2018
Earlier work this paper cites.
Hp-gan: Probabilistic 3d human motion prediction via gan
Barsoum, E., Kender, J., and Liu, Z. (2018) · 2018
Earlier work this paper cites.
Embodied question answering
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., and Batra, D. (2018) · 2018
Earlier work this paper cites.
Convolutional sequence to sequence model for human dynamics
Li, C., Zhang, Z., Lee, W. S., and Lee, G. H. (2018) · 2018
Earlier work this paper cites.
Human motion modeling using dvgans
Lin, X. and Amer, M. R. (2018) · 2018
Earlier work this paper cites.
Language2pose: Natural language grounded pose forecasting
Ahuja, C. and Morency, L.-P. (2019) · 2019
Earlier work this paper cites.
Holistic++ scene understanding: Single-view 3d holistic scene parsing and human pose estimation with human-object interaction and physical commonsense
Chen, Y., Huang, S., Yuan, T., Qi, S., Zhu, Y., and Zhu, S.-C. (2019) · 2019
Cited alongside, same era.
Splitnet: Sim2sim and task2task transfer for embodied visual navigation
Gordon, D., Kadian, A., Parikh, D., Hoffman, J., and Batra, D. (2019) · 2019
Cited alongside, same era.
Resolving 3d human pose ambiguities with 3d scene constraints
Hassan, M., Choutas, V., Tzionas, D., and Black, M. J. (2019) · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C. and Toutanova, L. K. (2019) · 2019
Cited alongside, same era.
Functional workspace optimization via learning personal preferences from virtual experiences
Liang, W., Liu, J., Lang, Y., Ning, B., and Yu, L.-F. (2019) · 2019
Cited alongside, same era.
Dlow: Diversifying latent flows for diverse human motion prediction
Yuan, Y. and Kitani, K. (2020) · 2020
Later among the works it cites.
Scanqa: 3d question answering for spatial scene understanding
Azuma, D., Miyanishi, T., Kurita, S., and Kawanabe, M. (2021) · 2021
Later among the works it cites.
Yourefit: Embodied reference understanding with language and gesture
Chen, Y., Li, Q., Kong, D., Kei, Y. L., Gao, T., Zhu, Y., and Huang, S. (2021) · 2021
Later among the works it cites.
Synthesis of compositional animations from textual descriptions
Ghosh, A., Cheema, N., Oguz, C., Theobalt, C., and Slusallek, P. (2021) · 2021
Later among the works it cites.
Vlgrammar: Grounded grammar induction of vision and language
Hong, Y., Li, Q., Zhu, S.-C., and Huang, S. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Amass: Archive of motion capture as surface shapes
Mahmood, N., Ghorbani, N., Troje, N. F., Pons-Moll, G., and Black, M. J. (2019) · 2019
Cited alongside, same era.
Expressive body capture: 3d hands, face, and body from a single image
Pavlakos, G., Choutas, V., Ghorbani, N., Bolkart, T., Osman, A. A., Tzionas, D., and Black, M. J. (2019) · 2019
Cited alongside, same era.
Neural state machine for character-scene interactions
Starke, S., Zhang, H., Komura, T., and Saito, J. (2019) · 2019
Cited alongside, same era.
On the continuity of rotation representations in neural networks
Zhou, Y., Barnes, C., Lu, J., Yang, J., and Li, H. (2019) · 2019
Cited alongside, same era.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Achlioptas, P., Abdelreheem, A., Xia, F., Elhoseiny, M., and Guibas, L. (2020) · 2020
Cited alongside, same era.
Long-term human motion prediction with scene context
Cao, Z., Gao, H., Mangalam, K., Cai, Q.-Z., Vo, M., and Malik, J. (2020) · 2020
Cited alongside, same era.
Scanrefer: 3d object localization in rgb-d scans using natural language
Chen, D. Z., Chang, A. X., and Nießner, M. (2020) · 2020
Cited alongside, same era.
Petrovich, M., Black, M. J., and Varol, G. (2021) · 2021
Later among the works it cites.
Babel: Bodies, action and behavior with english labels
Punnakkal, A. R., Chandrasekaran, A., Athanasiou, N., Quiros-Ramirez, A., and Black, M. J. (2021) · 2021
Later among the works it cites.
Physics-based human motion estimation and synthesis from videos
Xie, K., Wang, T., Iqbal, U., Guo, Y., Fidler, S., and Shkurti, F. (2021) · 2021
Later among the works it cites.
Ye, S., Chen, D., Han, S., and Liao, J. (2021) · 2021
Later among the works it cites.
We are more than our joints: Predicting how 3d bodies move
Zhang, Y., Black, M. J., and Tang, S. (2021) · 2021
Later among the works it cites.
Point transformer
Zhao, H., Jiang, L., Jia, J., Torr, P. H., and Koltun, V. (2021) · 2021
Later among the works it cites.
Capturing and inferring dense full-body human-scene contact
Huang, C.-H. P., Yi, H., Höschle, M., Safroshkin, M., Alexiadis, T., Polikovsky, S., Scharstein, D., and Black, M. J. (2022) · 2022
Closest in time.
Understanding embodied reference with touch-line transformer
Li, Y., Chen, X., Zhao, H., Gong, J., Zhou, G., Rossano, F., and Zhu, Y. (2022) · 2022
Closest in time.
Languagerefer: Spatial-language model for 3d visual grounding
Roh, J., Desingh, K., Farhadi, A., and Fox, D. (2022) · 2022
Closest in time.
Language grounding with 3d objects
Thomason, J., Shridhar, M., Bisk, Y., Paxton, C., and Zettlemoyer, L. (2022) · 2022
Closest in time.
Human-aware object placement for visual environment reconstruction
Yi, H., Huang, C.-H. P., Tzionas, D., Kocabas, M., Hassan, M., Tang, S., Thies, J., and Black, M. J. (2022) · 2022
Closest in time.