Fetching the paper…
Reading the bibliography…
We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - synthetic robot data generated from video world models.
Dynamic multi-path neural network
Y. Su, S. Zhou, Y. Wu, T. Su, D. Liang, J. Liu, D. Zheng, Y. Wang, J. Yan, and X. Hu · 2019
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al · 2022
Earlier work this paper cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos, 2022
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Earlier work this paper cites.
Cacti: A framework for scalable multi-task multi-scene visual imitation learning
Z. Mandi, H. Bharadhwaj, V. Moens, S. Song, A. Rajeswaran, and V. Kumar · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Earlier work this paper cites.
Dexmv: Imitation learning for dexterous manipulation from human videos
Y. Qin, Y.-H. Wu, S. Liu, H. Jiang, R. Yang, Y. Fu, and X. Wang · 2022
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Earlier work this paper cites.
Learning universal policies via text-guided video generation
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Earlier work this paper cites.
Mimicgen: A data generation system for scalable robot learning using human demonstrations
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y. Narang, L. Fan, Y. Zhu, and D. Fox · 2023
Earlier work this paper cites.
Imitating task and motion planning with visuomotor transformers
M. Dalal, A. Mandlekar, C. R. Garrett, A. Handa, R. Salakhutdinov, and D. Fox · 2023
Earlier work this paper cites.
Maniskill2: A unified benchmark for generalizable manipulation skills
J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, Y. Tang, S. Tao, X. Wei, Y. Yao, et al · 2023
Earlier work this paper cites.
Scaling up and distilling down: Language-guided robot skill acquisition
H. Ha, P. Florence, and S. Song · 2023
Earlier work this paper cites.
Scaling robot learning with semantically imagined experience
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, J. Peralta, B. Ichter, et al · 2023
Earlier work this paper cites.
Genaug: Retargeting behaviors to unseen situations via generative augmentation
Z. Chen, S. Kiami, A. Gupta, and V. Kumar · 2023
Earlier work this paper cites.
Unleashing large-scale video generative pre-training for visual robot manipulation
H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong · 2023
Earlier work this paper cites.
An unbiased look at datasets for visuo-motor pre-training
S. Dasari, M. K. Srirama, U. Jain, and A. Gupta · 2023
Earlier work this paper cites.
Affordances from human videos as a versatile representation for robotics
S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak · 2023
Earlier work this paper cites.
Deft: Dexterous fine-tuning for real-world hand policies
A. Kannan, K. Shaw, S. Bahl, P. Mannam, and D. Pathak · 2023
Earlier work this paper cites.
Videodex: Learning dexterity from internet videos
K. Shaw, S. Bahl, and D. Pathak · 2023
Earlier work this paper cites.
Any-point trajectory modeling for policy learning
C. Wen, X. Lin, J. So, K. Chen, Q. Dou, Y. Gao, and P. Abbeel · 2023
Earlier work this paper cites.
Mimicplay: Long-horizon imitation learning by watching human play
C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y. Zhu, and A. Anandkumar · 2023
Earlier work this paper cites.
Zero-shot robot manipulation from passive human videos
H. Bharadhwaj, A. Gupta, S. Tulsiani, and V. Kumar · 2023
Earlier work this paper cites.
Learning continuous grasping function with a dexterous hand from human demonstrations
J. Ye, J. Wang, B. Huang, Y. Qin, and X. Wang · 2023
Cited alongside, same era.
π \pi 0: A vision-language-action flow model for general robot control, 2024
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Cited alongside, same era.
Cogvideox: Text-to-video diffusion models with an expert transformer
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al · 2024
Cited alongside, same era.
Hunyuanvideo: A systematic framework for large video generative models
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al · 2024
Cited alongside, same era.
Vision-based manipulation from single human video with open-world object graphs
Y. Zhu, A. Lim, P. Stone, and Y. Zhu · 2024
Later among the works it cites.
Equibot: Sim (3)-equivariant diffusion policy for generalizable and data efficient learning
J. Yang, Z.-a. Cao, C. Deng, R. Antonova, S. Song, and J. Bohg · 2024
Later among the works it cites.
Genie: Generative interactive environments, 2024
J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y. Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, and T. Rocktäschel · 2024
Later among the works it cites.
Learning to act without actions
D. Schmidt and M. Jiang · 2024
Later among the works it cites.
Videocon: Robust video-language alignment via contrast captions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Lin, W. Liu, C. Chen, J. Lu, W. Hu, T.-J. Fu, J. Allardice, Z. Lai, L. Song, B. Zhang, et al · 2024
Cited alongside, same era.
Pandora: Towards general world model with natural language actions and video states
J. Xiang, G. Liu, Y. Gu, Q. Gao, Y. Ning, Y. Zha, Z. Feng, T. Tao, S. Hao, Y. Shi, et al · 2024
Cited alongside, same era.
Robodreamer: Learning compositional world models for robot imagination
S. Zhou, Y. Du, J. Chen, Y. Li, D.-Y. Yeung, and C. Gan · 2024
Cited alongside, same era.
Learning to act from actionless videos through dense correspondences
P.-C. Ko, J. Mao, Y. Du, S.-H. Sun, and J. B. Tenenbaum · 2024
Cited alongside, same era.
Learning interactive real-world simulators
S. Yang, Y. Du, S. K. S. Ghasemipour, J. Tompson, L. P. Kaelbling, D. Schuurmans, and P. Abbeel · 2024
Cited alongside, same era.
Video language planning
Y. Du, S. Yang, P. Florence, F. Xia, A. Wahid, brian ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, L. P. Kaelbling, A. Zeng, and J. Tompson · 2024
Cited alongside, same era.
Robocasa: Large-scale simulation of everyday tasks for generalist robots
S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y. Zhu · 2024
Cited alongside, same era.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Cited alongside, same era.
H. Bansal, Y. Bitton, I. Szpektor, K.-W. Chang, and A. Grover · 2024
Later among the works it cites.
Gemini robotics: Bringing ai into the physical world
G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, et al · 2025
Closest in time.
Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, S. Gao, X. He, X. Huang, S. Jiang, et al · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al · 2025
Closest in time.
Cosmos world foundation model platform for physical ai
N. Agarwal, A. Ali, M. Bala, Y. Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y. Chen, Y. Cui, Y. Ding, et al · 2025
Closest in time.
Wan: Open and advanced large-scale video generative models
A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, et al · 2025
Closest in time.
Latent action pretraining from videos
S. Ye, J. Jang, B. Jeon, S. J. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y.-W. Chao, B. Y. Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo · 2025
Closest in time.
Do generative video models learn physical principles from watching videos?
S. Motamed, L. Culp, K. Swersky, P. Jaini, and R. Geirhos · 2025
Closest in time.
Worldscore: A unified evaluation benchmark for world generation
H. Duan, H.-X. Yu, S. Chen, L. Fei-Fei, and J. Wu · 2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin · 2025
Closest in time.
Physics-driven data generation for contact-rich manipulation via trajectory optimization
L. Yang, H. Suh, T. Zhao, B. P. Graesdal, T. Kelestemur, J. Wang, T. Pang, and R. Tedrake · 2025
Closest in time.
Cosmos-transfer1: Conditional world generation with adaptive multimodal control
H. A. Alhaija, J. Alvarez, M. Bala, T. Cai, T. Cao, L. Cha, J. Chen, M. Chen, F. Ferroni, S. Fidler, et al · 2025
Closest in time.
Solving new tasks by adapting internet video knowledge
C. Luo, Z. Zeng, Y. Du, and C. Sun · 2025
Closest in time.
S. Li, Y. Gao, D. Sadigh, and S. Song · 2025
Closest in time.
Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets
C. Zhu, R. Yu, S. Feng, B. Burchfiel, P. Shah, and A. Gupta · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, et al · 2025
Closest in time.
Moto: Latent motion token as the bridging language for learning robot manipulation from videos, 2025
Y. Chen, Y. Ge, W. Tang, Y. Li, Y. Ge, M. Ding, Y. Shan, and X. Liu · 2025
Closest in time.
Videoworld: Exploring knowledge learning from unlabeled videos, 2025
Z. Ren, Y. Wei, X. Guo, Y. Zhao, B. Kang, J. Feng, and X. Jin · 2025
Closest in time.
Univla: Learning to act anywhere with task-centric latent actions
Q. Bu, Y. Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li · 2025
Closest in time.
Adaworld: Learning adaptable world models with latent actions
S. Gao, S. Zhou, Y. Du, J. Zhang, and C. Gan · 2025
Closest in time.
Lerobot: Making ai for robotics more accessible with end-to-end learning, 2024
R. Cadene, S. Alibert, A. Soare, Q. Gallouedec, and T. Wolf · 2025
Closest in time.