Fetching the paper…
Reading the bibliography…
Human pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
ORB-SLAM: A versatile and accurate monocular SLAM system
Raul Mur-Artal, Jose Maria Martinez Montiel, and Juan D Tardos · 2015
Earlier work this paper cites.
Body talk: Crowdshaping realistic 3D avatars with words
Stephan Streuber, M Alejandra Quiros-Ramirez, Matthew Q Hill, Carina A Hahn, Silvia Zuffi, Alice O’Toole, and Michael J Black · 2016
Earlier work this paper cites.
Monocular 3D human pose estimation in the wild using improved cnn supervision
Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt · 2017
Earlier work this paper cites.
ORB-SLAM: An open-source slam system for monocular, stereo, and RGB-D cameras
Raul Mur-Artal and Juan D Tardós · 2017
Earlier work this paper cites.
AI challenger: A large-scale dataset for going deeper in image understanding
Jiahong Wu, He Zheng, Bo Zhao, Yixin Li, Baoming Yan, Rui Liang, Wenjia Wang, Shipei Zhou, Guosen Lin, Yanwei Fu, et al · 2017
Earlier work this paper cites.
End-to-end recovery of human shape and pose
Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik · 2018
Earlier work this paper cites.
Recovering accurate 3d human pose in the wild using imus and a moving camera
Timo Von Marcard, Roberto Henschel, Michael J Black, Bodo Rosenhahn, and Gerard Pons-Moll · 2018
Earlier work this paper cites.
Centernet: Keypoint triplets for object detection
Kaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian · 2019
Earlier work this paper cites.
CAM-Convs: Camera-aware multi-scale convolutions for single-view depth
Jose M. Facil, Benjamin Ummenhofer, Huizhong Zhou, Luis Montesano, Thomas Brox, and Javier Civera · 2019
Earlier work this paper cites.
Learning 3D human dynamics from video
Angjoo Kanazawa, Jason Y Zhang, Panna Felsen, and Jitendra Malik · 2019
Earlier work this paper cites.
Learning to reconstruct 3D human pose and shape via model-fitting in the loop
Nikos Kolotouros, Georgios Pavlakos, Michael J Black, and Kostas Daniilidis · 2019
Earlier work this paper cites.
Expressive body capture: 3D hands, face, and body from a single image
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black · 2019
Earlier work this paper cites.
Human mesh recovery from monocular images via a skeleton-disentangled representation
Yu Sun, Yun Ye, Wu Liu, Wenpeng Gao, Yili Fu, and Tao Mei · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Earlier work this paper cites.
Pose2Mesh: Graph convolutional network for 3D human pose and mesh recovery from a 2D human pose
Hongsuk Choi, Gyeongsik Moon, and Kyoung Mu Lee · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Coherent reconstruction of multiple humans from a single image
Wen Jiang, Nikos Kolotouros, Georgios Pavlakos, Xiaowei Zhou, and Kostas Daniilidis · 2020
Earlier work this paper cites.
VIBE: Video inference for human body pose and shape estimation
Muhammed Kocabas, Nikos Athanasiou, and Michael J Black · 2020
Earlier work this paper cites.
3d human motion estimation via motion compression and refinement
Zhengyi Luo, S. Alireza Golestaneh, and Kris M. Kitani · 2020
Earlier work this paper cites.
I2L-MeshNet: Image-to-lixel prediction network for accurate 3d human pose and mesh estimation from a single RGB image
Gyeongsik Moon and Kyoung Mu Lee · 2020
Earlier work this paper cites.
Beyond static features for temporally consistent 3D human pose and shape from a video
Hongsuk Choi, Gyeongsik Moon, Ju Yong Chang, and Kyoung Mu Lee · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Mesh graphormer
Kevin Lin, Lijuan Wang, and Zicheng Liu · 2021
Earlier work this paper cites.
AGORA: Avatars in geography optimized for regression analysis
Priyanka Patel, Chun-Hao P Huang, Joachim Tesch, David T Hoffmann, Shashank Tripathi, and Michael J Black · 2021
Cited alongside, same era.
Action-conditioned 3D human motion synthesis with transformer VAE
Mathis Petrovich, Michael J Black, and Gül Varol · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
HUMOR: 3D human motion model for robust pose estimation
Davis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang, Srinath Sridhar, and Leonidas J Guibas · 2021
Cited alongside, same era.
Monocular, one-stage, regression of multiple 3D people
Yu Sun, Qian Bao, Wu Liu, Yili Fu, Michael J Black, and Tao Mei · 2021
Cited alongside, same era.
DINOv2: Learning robust visual features without supervision, 2023
Maxime Oquab, Timothée Darcet, Theo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicolas Ballas, Gabriel Synnaeve, Ishan Misra, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2023
Later among the works it cites.
WHAM: Reconstructing world-grounded humans with accurate 3D motion
Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J Black · 2023
Later among the works it cites.
TRACE: 5D temporal regression of avatars with dynamic cameras in 3D environments
Yu Sun, Qian Bao, Wu Liu, Tao Mei, and Michael J Black · 2023
Later among the works it cites.
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano · 2023
Later among the works it cites.
ReFit: Recurrent fitting network for 3D human recovery
Yufu Wang and Kostas Daniilidis · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zachary Teed and Jia Deng · 2021
Cited alongside, same era.
PyMAF: 3D human pose and shape regression with pyramidal mesh alignment feedback loop
Hongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang, Yebin Liu, Limin Wang, and Zhenan Sun · 2021
Cited alongside, same era.
Accurate 3D body shape regression using metric and semantic attributes
Vasileios Choutas, Lea Müller, Chun-Hao P. Huang, Siyu Tang, Dimitrios Tzionas, and Michael J. Black · 2022
Cited alongside, same era.
BodySLAM: joint camera localisation, mapping, and human motion tracking
Dorian F Henning, Tristan Laidlow, and Stefan Leutenegger · 2022
Cited alongside, same era.
Capturing and inferring dense full-body human-scene contact
Chun-Hao P Huang, Hongwei Yi, Markus Höschle, Matvey Safroshkin, Tsvetelina Alexiadis, Senya Polikovsky, Daniel Scharstein, and Michael J Black · 2022
Cited alongside, same era.
CLIFF: Carrying location information in full frames into human pose and shape estimation
Zhihao Li, Jianzhuang Liu, Zhensong Zhang, Songcen Xu, and Youliang Yan · 2022
Cited alongside, same era.
Posegpt: Quantization-based 3d human motion generation and forecasting
Thomas Lucas, Fabien Baradel, Philippe Weinzaepfel, and Grégory Rogez · 2022
Cited alongside, same era.
Later among the works it cites.
Hu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer · 2023
Later among the works it cites.
Decoupling human and camera motion from videos in the wild
Vickie Ye, Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa · 2023
Later among the works it cites.
Hi4D: 4D instance segmentation of close human interaction
Yifei Yin, Chen Guo, Manuel Kaufmann, Juan Jose Zarate, Jie Song, and Otmar Hilliges · 2023
Later among the works it cites.
MotionFix: Text-driven 3D human motion editing
Nikos Athanasiou, Alpár Ceske, Markos Diomataris, Michael J. Black, and Gül Varol · 2024
Later among the works it cites.
Multi-HMR: Multi-person whole-body human mesh recovery in a single shot
Fabien Baradel, Matthieu Armando, Salma Galaaoui, Romain Brégier, Philippe Weinzaepfel, Grégory Rogez, and Thomas Lucas · 2024
Later among the works it cites.
PoseEmbroider: Towards a 3D, visual, semantic-aware human pose representation
Ginger Delmas, Philippe Weinzaepfel, Francesc Moreno-Noguer, and Grégory Rogez · 2024
Later among the works it cites.
TokenHMR: Advancing human mesh recovery with a tokenized pose representation
Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, Yao Feng, and Michael J Black · 2024
Later among the works it cites.
ChatPose: Chatting about 3D human pose
Yao Feng, Jing Lin, Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, and Michael J. Black · 2024
Later among the works it cites.
PACE: Human and camera motion estimation from in-the-wild videos
Muhammed Kocabas, Ye Yuan, Pavlo Molchanov, Yunrong Guo, Michael J Black, Otmar Hilliges, Jan Kautz, and Umar Iqbal · 2024
Later among the works it cites.
Generative proxemics: A prior for 3D social interaction from images
Lea Müller, Vickie Ye, Georgios Pavlakos, Michael J. Black, and Angjoo Kanazawa · 2024
Later among the works it cites.
BodyShapeGPT: SMPL body shape manipulation with LLMs
Baldomero R. Árbol and Dan Casas · 2024
Later among the works it cites.
World-grounded human motion recovery via gravity-view coordinates
Zehong Shen, Huaijin Pi, Yan Xia, Zhi Cen, Sida Peng, Zechen Hu, Hujun Bao, Ruizhen Hu, and Xiaowei Zhou · 2024
Later among the works it cites.
Pose priors from language models
Sanjay Subramanian, Evonne Ng, Lea Müller, Dan Klein, Shiry Ginosar, and Trevor Darrell · 2024
Later among the works it cites.
AiOS: All-in-one-stage expressive human pose and shape estimation
Qingping Sun, Yanjun Wang, Ailing Zeng, Wanqi Yin, Chen Wei, Wenjia Wang, Haiyi Mei, Chi-Sing Leung, Ziwei Liu, Lei Yang, and Zhongang Cai · 2024
Later among the works it cites.
Deep patch visual odometry
Zachary Teed, Lahav Lipson, and Jia Deng · 2024
Later among the works it cites.
TRAM: Global trajectory and motion of 3d humans from in-the-wild videos
Yufu Wang, Ziyun Wang, Lingjie Liu, and Kostas Daniilidis · 2024
Later among the works it cites.
Synergistic global-space camera and human reconstruction from videos
Yizhou Zhao, Tuanfeng Yang Wang, Bhiksha Raj, Min Xu, Jimei Yang, and Chun-Hao Paul Huang · 2024
Later among the works it cites.
Sapiens: Foundation for human vision models
Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez, Su Zhaoen, Austin James, Peter Selednik, Stuart Anderson, and Shunsuke Saito · 2025
Closest in time.
Camerahmr: Aligning people with perspective
Priyanka Patel and Michael J. Black · 2025
Closest in time.