Fetching the paper…
Reading the bibliography…
Recent robot learning methods commonly rely on imitation learning from massive robotic dataset collected with teleoperation.
Avid: Learning multi-stage tasks via pixel-level translation of human videos
L. Smith, N. Dhawan, M. Zhang, P. Abbeel, and S. Levine · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. N. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Earlier work this paper cites.
Prototypical contrastive learning of unsupervised representations
J. Li, P. Zhou, C. Xiong, and S. C. Hoi · 2020
Earlier work this paper cites.
Learning by watching: Physical imitation of manipulation skills from human videos
H. Xiong, Q. Li, Y.-C. Chen, H. Bharadhwaj, S. Sinha, and A. Garg · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. S. Ryoo, G. Salazar, P. R. Sanketi, K. Sayed, J. Singh, S. A. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. H. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Earlier work this paper cites.
Dexmv: Imitation learning for dexterous manipulation from human videos
Y. Qin, Y.-H. Wu, S. Liu, H. Jiang, R. Yang, Y. Fu, and X. Wang · 2022
Earlier work this paper cites.
Siamese prototypical contrastive learning
S. Mo, Z. Sun, and C. Li · 2022
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, K. Choromanski, T. Ding, D. Driess, K. A. Dubey, C. Finn, P. R. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y. Kuang, I. Leal, S. Levine, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. S. Ryoo, G. Salazar, P. R. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. H. Vuong, A. Wahid, S. Welker, P. Wohlhart, T. Xiao, T. Yu, and B. Zitkovich · 2023
Earlier work this paper cites.
Deft: Dexterous fine-tuning for real-world hand policies
A. Kannan, K. Shaw, S. Bahl, P. Mannam, and D. Pathak · 2023
Earlier work this paper cites.
Xskill: Cross embodiment skill discovery
M. Xu, Z. Xu, C. Chi, M. Veloso, and S. Song · 2023
Earlier work this paper cites.
Mimicplay: Long-horizon imitation learning by watching human play
C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y. Zhu, and A. Anandkumar · 2023
Earlier work this paper cites.
Any-point trajectory modeling for policy learning
C. Wen, X. Lin, J. So, K. Chen, Q. Dou, Y. Gao, and P. Abbeel · 2023
Earlier work this paper cites.
Rt-trajectory: Robotic task generalization via hindsight trajectory sketches
J. Gu, S. Kirmani, P. Wohlhart, Y. Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, et al · 2023
Earlier work this paper cites.
Learning universal policies via text-guided video generation
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel · 2023
Cited alongside, same era.
One-shot visual imitation via attributed waypoints and demonstration augmentation
M. Chang and S. Gupta · 2023
Cited alongside, same era.
Robotube: Learning household manipulation from human videos with simulated twin environments
H. Xiong, H. Fu, J. Zhang, C. Bao, Q. Zhang, Y. Huang, W. Xu, A. Garg, and C. Lu · 2023
Cited alongside, same era.
Learning video-conditioned policies for unseen manipulation tasks
E. Chane-Sane, C. Schmid, and I. Laptev · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Cited alongside, same era.
Latent action pretraining from videos
S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y.-W. Chao, B. Y. Lin, et al · 2024
Later among the works it cites.
Dreamitate: Real-world visuomotor policy learning via video generation
J. Liang, R. Liu, E. Ozguroglu, S. Sudhakar, A. Dave, P. Tokmakov, S. Song, and C. Vondrick · 2024
Later among the works it cites.
Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation
H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani · 2024
Later among the works it cites.
Generative image as action models
M. Shridhar, Y. L. Lo, and S. James · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Pearce, T. Rashid, A. Kanervisto, D. Bignell, M. Sun, R. Georgescu, S. V. Macua, S. Z. Tan, I. Momennejad, K. Hofmann, and S. Devlin · 2023
Cited alongside, same era.
Goal-conditioned imitation learning using score-based diffusion policies
M. Reuss, M. X. Li, X. Jia, and R. Lioutikov · 2023
Cited alongside, same era.
Anyteleop: A general vision-based dexterous robot arm-hand teleoperation system
Y. Qin, W. Yang, B. Huang, K. V. Wyk, H. Su, X. Wang, Y.-W. Chao, and D. Fox · 2023
Cited alongside, same era.
Egomimic: Scaling imitation learning via egocentric video
S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu · 2024
Cited alongside, same era.
Vividex: Learning vision-based dexterous manipulation from human videos
Z. Chen, S. Chen, E. Arlaud, I. Laptev, and C. Schmid · 2024
Cited alongside, same era.
Dexcap: Scalable and portable mocap data collection system for dexterous manipulation
C. Wang, H. Shi, W. Wang, R. Zhang, L. Fei-Fei, and C. K. Liu · 2024
Cited alongside, same era.
Vid2robot: End-to-end video-conditioned policy learning with cross-attention transformers
V. Jain, M. Attarian, N. J. Joshi, A. Wahid, D. Driess, Q. Vuong, P. R. Sanketi, P. Sermanet, S. Welker, C. Chan, et al · 2024
Cited alongside, same era.
Z. Qian, M. You, H. Zhou, X. Xu, H. Fu, J. Xue, and B. He · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. R. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. L. Tan, P. R. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine · 2024
Later among the works it cites.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
M. Reuss, Ö. E. Yagmurlu, F. Wenzel, and R. Lioutikov · 2024
Later among the works it cites.
Rdt-1b: a diffusion foundation model for bimanual manipulation
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu · 2024
Later among the works it cites.
Diffusion actor-critic with entropy regulator
Y. Wang, L. Wang, Y. Jiang, W. Zou, T. Liu, X. Song, W. Wang, L. Xiao, J. Wu, J. Duan, et al · 2024
Later among the works it cites.
Sampling from energy-based policies using diffusion
V. Jain, T. Akhound-Sadegh, and S. Ravanbakhsh · 2024
Later among the works it cites.
Learning multimodal behaviors from scratch with diffusion policy gradient
S. Li, R. Krohn, T. Chen, A. Ajay, P. Agrawal, and G. Chalvatzaki · 2024
Later among the works it cites.
Entropy-regularized diffusion policy with q-ensembles for offline reinforcement learning
R. Zhang, Z. Luo, J. Sjölund, T. Schön, and P. Mattsson · 2024
Later among the works it cites.
Wilor: End-to-end 3d hand localization and reconstruction in-the-wild
R. A. Potamias, J. Zhang, J. Deng, and S. Zafeiriou · 2024
Later among the works it cites.