Fetching the paper…
Reading the bibliography…
Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms.
Discriminative and adaptive imitation in uni-manual and bi-manual tasks
A. G. Billard, S. Calinon, and F. Guenter · 2006
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2015
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Sim-to-real transfer of robotic control with dynamics randomization
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
Closing the sim-to-real loop: Adapting simulation randomization with real world experience
Y. Chebotar, A. Handa, V. Makoviychuk, M. Macklin, J. Issac, N. Ratliff, and D. Fox · 2019
Earlier work this paper cites.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg · 2020
Earlier work this paper cites.
Sim2real2sim: Bridging the gap between simulation and real-world in flexible object manipulation
P. Chang and T. Padif · 2020
Earlier work this paper cites.
Xirl: Cross-embodiment inverse reinforcement learning
K. Zakka, A. Zeng, P. R. Florence, J. Tompson, J. Bohg, and D. Dwibedi · 2021
Earlier work this paper cites.
Maniskill: Generalizable manipulation skill benchmark with large-scale demonstrations
T. Mu, Z. Ling, F. Xiang, D. Yang, X. Li, S. Tao, Z. Huang, Z. Jia, and H. Su · 2021
Earlier work this paper cites.
Planar robot casting with real2sim2real self-supervised learning
V. Lim, H. Huang, L. Y. Chen, J. Wang, J. Ichnowski, D. Seita, M. Laskey, and K. Goldberg · 2021
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Earlier work this paper cites.
Robotic telekinesis: Learning a robotic hand imitator by watching humans on youtube
A. Sivakumar, K. Shaw, and D. Pathak · 2022
Earlier work this paper cites.
Videodex: Learning dexterity from internet videos
K. Shaw, S. Bahl, and D. Pathak · 2022
Earlier work this paper cites.
Dexterous imitation made easy: A learning-based framework for efficient dexterous manipulation
S. P. Arunachalam, S. Silwal, B. Evans, and L. Pinto · 2022
Earlier work this paper cites.
Learning continuous grasping function with a dexterous hand from human demonstrations
J. Ye, J. Wang, B. Huang, Y. Qin, and X. Wang · 2022
Earlier work this paper cites.
Human-to-robot imitation in the wild
S. Bahl, A. Gupta, and D. Pathak · 2022
Earlier work this paper cites.
Learning to imitate object interactions from internet videos
A. Patel, A. Wang, I. Radosavovic, and J. Malik · 2022
Earlier work this paper cites.
Graph inverse reinforcement learning from diverse videos
S. Kumar, J. Zamora, N. Hansen, R. Jangir, and X. Wang · 2022
Earlier work this paper cites.
Behavior-1k: A benchmark for embodied ai with 1, 000 everyday activities and realistic simulation
C. Li, R. Zhang, J. Wong, C. Gokmen, S. Srivastava, R. Martín-Martín, C. Wang, G. Levine, M. Lingelbach, J. Sun, M. Anvari, M. Hwang, M. Sharma, A. Aydin, D. Bansal, S. Hunter, K.-Y. Kim, A. Lou, C. R. Matthews, I. Villa-Renteria, J. H. Tang, C. Tang, F. Xia, S. Savarese, H. Gweon, C. K. Liu, J. Wu, and L. Fei-Fei · 2022
Earlier work this paper cites.
Open x-embodiment: Robotic learning datasets and rt-x models
A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Cited alongside, same era.
Zero-shot robot manipulation from passive human videos
H. Bharadhwaj, A. Gupta, S. Tulsiani, and V. Kumar · 2023
Cited alongside, same era.
Mimicplay: Long-horizon imitation learning by watching human play
Okami: Teaching humanoid robots manipulation skills through single video imitation
J. Li, Y. Zhu, Y. Xie, Z. Jiang, M. Seo, G. Pavlakos, and Y. Zhu · 2024
Later among the works it cites.
Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation
H. Bharadhwaj, R. Mottaghi, A. Gupta, and S. Tulsiani · 2024
Later among the works it cites.
Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation
H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani · 2024
Later among the works it cites.
One-shot imitation under mismatched execution
K. Kedia, P. Dan, A. Chao, M. A. Pace, and S. Choudhury · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y. Zhu, and A. Anandkumar · 2023
Cited alongside, same era.
Foundationpose: Unified 6d pose estimation and tracking of novel objects
B. Wen, W. Yang, J. Kautz, and S. T. Birchfield · 2023
Cited alongside, same era.
Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators
P. Wu, Y. Shentu, Z. Yi, X. Lin, and P. Abbeel · 2023
Cited alongside, same era.
One-shot imitation learning: A pose estimation perspective
P. Vitiello, K. Dreczkowski, and E. Johns · 2023
Cited alongside, same era.
Towards generalizable zero-shot manipulation via translating human interaction plans
H. Bharadhwaj, A. Gupta, V. Kumar, and S. Tulsiani · 2023
Cited alongside, same era.
XSkill: Cross embodiment skill discovery
M. Xu, Z. Xu, C. Chi, M. Veloso, and S. Song · 2023
Cited alongside, same era.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. B. Girshick · 2023
Cited alongside, same era.
Ditto in the house: Building articulation models of indoor scenes through interactive perception
C.-C. Hsu, Z. Jiang, and Y. Zhu · 2023
Cited alongside, same era.
I. Güzey, Y. Dai, G. Savva, R. M. Bhirangi, and L. Pinto · 2024
Later among the works it cites.
Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation
M. M. L. Torné, A. Simeonov, Z. Li, A. Chan, T. Chen, A. Gupta, and P. Agrawal · 2024
Later among the works it cites.
From imitation to refinement – residual rl for precise assembly
L. L. Ankile, A. Simeonov, I. Shenfeld, M. M. L. Torné, and P. Agrawal · 2024
Later among the works it cites.
Robot see robot do: Imitating articulated object manipulation with monocular 4d reconstruction
J. Kerr, C. M. Kim, M. Wu, B. Yi, Q. Wang, K. Goldberg, and A. Kanazawa · 2024
Later among the works it cites.
J. Xu, W. Cheng, Y. Gao, X. Wang, S. Gao, and Y. Shan · 2024
Later among the works it cites.
Automated creation of digital cousins for robust policy learning
T. Dai, J. Wong, Y. Jiang, C. Wang, C. Gokmen, R. Zhang, J. Wu, and F.-F. Li · 2024
Later among the works it cites.
π 0.5 \pi\;0.5 : A vision-language-action model with open-world generalization
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. R. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky · 2025
Closest in time.
Motion tracks: A unified representation for human-robot transfer in few-shot imitation learning
J. Ren, P. Sundaresan, D. Sadigh, S. Choudhury, and J. Bohg · 2025
Closest in time.
A real-to-sim-to-real approach to robotic manipulation with vlm-generated iterative keypoint rewards
S. Patel, X. Yin, W. Huang, S. Garg, H. Nayyeri, F.-F. Li, S. Lazebnik, and Y. Li · 2025
Closest in time.
Video2policy: Scaling up manipulation tasks in simulation through internet videos
W. Ye, F. Liu, Z. Ding, Y. Gao, O. Rybkin, and P. Abbeel · 2025
Closest in time.
Crossing the human-robot embodiment gap with sim-to-real rl using one human demonstration
T. Ga, W. Lum, O. Y. Lee, C. K. Liu, J. Bohg, and P.-M. H. Pose · 2025
Closest in time.
Polycam, 2020
Polycam · 2025
Closest in time.
Drawer: Digital reconstruction and articulation with environment realism
H. Xia, E. Su, M. Memmel, A. Jain, R. Yu, N. Mbiziwo-Tiapo, A. Farhadi, A. Gupta, S. Wang, and W.-C. Ma · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
Nvidia, J. Bjorck, F. Castaneda, N. Cherniadev, X. Da, R. Ding, LinxiJimFan, Y. Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y. L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y. Xie, Y. Xu, Z.-T. Xu, S. Ye, Z. Yu, A. Zhang, H. Zhang, Y. Zhao, R. Zheng, and Y. Zhu · 2025
Closest in time.
Drawer: Digital reconstruction and articulation with environment realism
H. Xia, E. Su, M. Memmel, A. Jain, R. Yu, N. Mbiziwo-Tiapo, A. Farhadi, A. Gupta, S. Wang, and W.-C. Ma · 2025
Closest in time.
St4rtrack: Simultaneous 4d reconstruction and tracking in the world
H. Feng, J. Zhang, Q. Wang, Y. Ye, P. Yu, M. J. Black, T. Darrell, and A. Kanazawa · 2025
Closest in time.