Fetching the paper…
Reading the bibliography…
How can we teach humanoids to climb staircases and sit on chairs using the surrounding environment context? Arguably, the simplest way is to just show them-casually capture a human motion video and feed it to humanoids.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 1901
Earlier work this paper cites.
No-regret reductions for imitation learning and structured prediction
S. Ross, G. J. Gordon, and J. A. Bagnell · 2010
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
DeepMimic
X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne · 2018
Earlier work this paper cites.
SFV: reinforcement learning of physical skills from videos
X. B. Peng, A. Kanazawa, J. Malik, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Per-contact iteration method for solving contact dynamics
J. Hwangbo, J. Lee, and M. Hutter · 2018
Earlier work this paper cites.
End-to-end recovery of human shape and pose
A. Kanazawa, M. J. Black, D. W. Jacobs, and J. Malik · 2018
Earlier work this paper cites.
2d/3d pose estimation and action recognition using multitask deep learning
D. C. Luvizon, D. Picard, and H. Tabia · 2018
Earlier work this paper cites.
Probabilistic terrain mapping for mobile robots with uncertain localization
P. Fankhauser, M. Bloesch, and M. Hutter · 2018
Earlier work this paper cites.
Learning quadrupedal locomotion over challenging terrain
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2020
Earlier work this paper cites.
Vibe: Video inference for human body pose and shape estimation
M. Kocabas, N. Athanasiou, and M. J. Black · 2020
Earlier work this paper cites.
Robust motion in-betweening
F. G. Harvey, M. Yurick, D. Nowrouzezahrai, and C. Pal · 2020
Earlier work this paper cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State · 2021
Earlier work this paper cites.
RMA: rapid motor adaptation for legged robots
A. Kumar, Z. Fu, D. Pathak, and J. Malik · 2021
Earlier work this paper cites.
Human dynamics from monocular video with dynamic camera movements
R. Yu, H. Park, and J. Lee · 2021
Earlier work this paper cites.
Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans
S. Peng, Y. Zhang, Y. Xu, Q. Wang, Q. Shuai, H. Bao, and X. Zhou · 2021
Earlier work this paper cites.
Simpoe: Simulated character control for 3d human pose estimation
Y. Yuan, S.-E. Wei, T. Simon, K. Kitani, and J. Saragih · 2021
Earlier work this paper cites.
Learning motion priors for 4d human body capture in 3d scenes
S. Zhang, Y. Zhang, F. Bogo, M. Pollefeys, and S. Tang · 2021
Earlier work this paper cites.
Fast-lio2: Fast direct lidar-inertial odometry, 2021
W. Xu, Y. Cai, D. He, J. Lin, and F. Zhang · 2021
Earlier work this paper cites.
rl-games: A high-performance framework for reinforcement learning
D. Makoviichuk and V. Makoviychuk · 2021
Earlier work this paper cites.
Legged locomotion in challenging terrains using egocentric vision, 2022
A. Agarwal, A. Kumar, J. Malik, and D. Pathak · 2022
Earlier work this paper cites.
Tracking people by predicting 3d appearance, location and pose
J. Rajasegaran, G. Pavlakos, A. Kanazawa, and J. Malik · 2022
Earlier work this paper cites.
Glamr: Global occlusion-aware human mesh recovery with dynamic cameras
Y. Yuan, U. Iqbal, P. Molchanov, K. Kitani, and J. Kautz · 2022
Earlier work this paper cites.
D &d: Learning human dynamics from dynamic camera
J. Li, S. Bian, C. Xu, G. Liu, G. Yu, and C. Lu · 2022
Cited alongside, same era.
Vitpose: Simple vision transformer baselines for human pose estimation
Y. Xu, J. Zhang, Q. Zhang, and D. Tao · 2022
Cited alongside, same era.
Capturing and inferring dense full-body human-scene contact
C.-H. P. Huang, H. Yi, M. Höschle, M. Safroshkin, T. Alexiadis, S. Polikovsky, D. Scharstein, and M. J. Black · 2022
Cited alongside, same era.
Learning to walk in minutes using massively parallel deep reinforcement learning
N. Rudin, D. Hoeller, P. Reist, and M. Hutter · 2022
Cited alongside, same era.
Real-world humanoid locomotion with reinforcement learning, 2023
I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath · 2023
Cited alongside, same era.
Multiphys: multi-person physics-aware 3d motion estimation
N. Ugrinovic, B. Pan, G. Pavlakos, D. Paschalidou, B. Shen, J. Sanchez-Riera, F. Moreno-Noguer, and L. Guibas · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V. Alwala, N. Carion, C.-Y. Wu, R. Girshick, P. Dollár, and C. Feichtenhofer · 2024
Later among the works it cites.
Grounded sam: Assembling open-world models for diverse visual tasks, 2024
T. Ren, S. Liu, A. Zeng, J. Lin, K. Li, H. Cao, J. Chen, X. Huang, Y. Chen, F. Yan, Z. Zeng, H. Zhang, F. Li, J. Yang, H. Li, Q. Jiang, and L. Zhang · 2024
Later among the works it cites.
Tram: Global trajectory and motion of 3d humans from in-the-wild videos
Y. Wang, Z. Wang, L. Liu, and K. Daniilidis · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Hoeller, N. Rudin, D. Sako, and M. Hutter · 2023
Cited alongside, same era.
Synthesizing physical character-scene interactions
M. Hassan, Y. Guo, T. Wang, M. Black, S. Fidler, and X. B. Peng · 2023
Cited alongside, same era.
Smpl: A skinned multi-person linear model
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black · 2023
Cited alongside, same era.
On the benefits of 3d pose and tracking for human action recognition
J. Rajasegaran, G. Pavlakos, A. Kanazawa, C. Feichtenhofer, and J. Malik · 2023
Cited alongside, same era.
Decoupling human and camera motion from videos in the wild
V. Ye, G. Pavlakos, J. Malik, and A. Kanazawa · 2023
Cited alongside, same era.
Learning physically simulated tennis skills from broadcast videos
Y. Yuan, V. Makoviychuk, Y. Guo, S. Fidler, X. Peng, and K. Fatahalian · 2023
Cited alongside, same era.
Neural kernel surface reconstruction
J. Huang, Z. Gojcic, M. Atzmon, O. Litany, S. Fidler, and F. Williams · 2023
Cited alongside, same era.
B. Yi, V. Ye, M. Zheng, Y. Li, L. Müller, G. Pavlakos, Y. Ma, J. Malik, and A. Kanazawa · 2024
Later among the works it cites.
Geocalib: Learning single-image calibration with geometric optimization
A. Veicht, P.-E. Sarlin, P. Lindenberger, and M. Pollefeys · 2024
Later among the works it cites.
Maskedmimic: Unified physics-based character control through masked motion inpainting
C. Tessler, Y. Guo, O. Nabati, G. Chechik, and X. B. Peng · 2024
Later among the works it cites.
Hand-object interaction pretraining from videos, 2024
H. G. Singh, A. Loquercio, C. Sferrazza, J. Wu, H. Qi, P. Abbeel, and J. Malik · 2024
Later among the works it cites.
Dextrah-g: Pixels-to-action dexterous arm-hand grasping with geometric fabrics, 2024
T. G. W. Lum, M. Matak, V. Makoviychuk, A. Handa, A. Allshire, T. Hermans, N. D. Ratliff, and K. V. Wyk · 2024
Later among the works it cites.
WHAM: Reconstructing world-grounded humans with accurate 3D motion
S. Shin, J. Kim, E. Halilaj, and M. J. Black · 2024
Later among the works it cites.
Align3r: Aligned monocular depth estimation for dynamic videos
J. Lu, T. Huang, P. Li, Z. Dou, C. Lin, Z. Cui, Z. Dong, S.-K. Yeung, W. Wang, and Y. Liu · 2024
Later among the works it cites.
Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoperation and learning, 2024
T. He, Z. Luo, X. He, W. Xiao, C. Zhang, W. Zhang, K. Kitani, C. Liu, and G. Shi · 2024
Later among the works it cites.
Asap: Aligning simulation and real-world physics for learning agile humanoid whole-body skills, 2025
T. He, J. Gao, W. Xiao, Y. Zhang, Z. Wang, J. Wang, Z. Luo, G. He, N. Sobanbab, C. Pan, Z. Yi, G. Qu, K. Kitani, J. Hodgins, L. J. Fan, Y. Zhu, C. Liu, and G. Shi · 2025
Closest in time.
Exbody2: Advanced expressive humanoid whole-body control, 2025
M. Ji, X. Peng, F. Liu, J. Li, G. Yang, X. Cheng, and X. Wang · 2025
Closest in time.
Beamdojo: Learning agile humanoid locomotion on sparse footholds
H. Wang, Z. Wang, J. Ren, Q. Ben, T. Huang, W. Zhang, and J. Pang · 2025
Closest in time.
Tokenhsi: Unified synthesis of physical human-scene interactions through task tokenization, 2025
L. Pan, Z. Yang, Z. Dou, W. Wang, B. Huang, B. Dai, T. Komura, and J. Wang · 2025
Closest in time.
Continuous 3d perception model with persistent state
Q. Wang, Y. Zhang, A. Holynski, A. A. Efros, and A. Kanazawa · 2025
Closest in time.
Joint optimization for 4d human-scene reconstruction in the wild
Z. Liu, J. Lin, W. Wu, and B. Zhou · 2025
Closest in time.
Viser: Imperative, web-based 3d visualization in python
B. Yi, C. M. Kim, J. Kerr, G. Wu, R. Feng, A. Zhang, J. Kulhanek, H. Choi, Y. Ma, M. Tancik, et al · 2025
Closest in time.
K. Zakka, B. Tabanpour, Q. Liao, M. Haiderbhai, S. Holt, J. Y. Luo, A. Allshire, E. Frey, K. Sreenath, L. A. Kahrs, C. Sferrazza, Y. Tassa, and P. Abbeel · 2025
Closest in time.
Pyroki: A modular toolkit for robot kinematic optimization, 2025
C. M. Kim, B. Yi, H. Choi, Y. Ma, K. Goldberg, and A. Kanazawa · 2025
Closest in time.
Clone: Closed-loop whole-body humanoid teleoperation for long-horizon tasks, 2025
Y. Li, Y. Lin, J. Cui, T. Liu, W. Liang, Y. Zhu, and S. Huang · 2025
Closest in time.