Fetching the paper…
Reading the bibliography…
We introduce Genie Envisioner (GE), a unified world foundation platform for robotic manipulation that integrates policy learning, evaluation, and simulation within a single video-generative framework.
An adaptive network that constructs and uses and internal model of its world
R. S. Sutton and A. G. Barto · 1981
Earlier work this paper cites.
Position referencing and consistent world modeling for mobile robots
R. Chatila and J.-P. Laumond · 1985
Earlier work this paper cites.
Mechanics of robotic manipulation
M. T. Mason · 2001
Earlier work this paper cites.
Task constrained motion planning in robot joint space
M. Stilman · 2007
Earlier work this paper cites.
Manipulation planning on constraint manifolds
D. Berenson, S. S. Srinivasa, D. Ferguson, and J. J. Kuffner · 2009
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
A mathematical introduction to robotic manipulation
R. M. Murray, Z. Li, and S. S. Sastry · 2017
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
F. Ebert, C. Finn, S. Dasari, A. Xie, A. Lee, and S. Levine · 2018
Earlier work this paper cites.
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Earlier work this paper cites.
General evaluation for instruction conditioned navigation using dynamic time warping
G. Ilharco, V. Jain, A. Ku, E. Ie, and J. Baldridge · 2019
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Deep dynamics models for learning dexterous manipulation
A. Nagabandi, K. Konolige, S. Levine, and V. Kumar · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al · 2022
Earlier work this paper cites.
Video diffusion models
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, et al · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Cited alongside, same era.
Hailuo AI
MiniMax · 2024
Later among the works it cites.
Robocasa: Large-scale simulation of everyday tasks for generalist robots
S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y. Zhu · 2024
Later among the works it cites.
Sora, 2024
OpenAI · 2024
Later among the works it cites.
T2v-compbench: A comprehensive benchmark for compositional text-to-video generation
K. Sun, K. Huang, X. Liu, Y. Wu, Z. Xu, Z. Li, and X. Liu · 2024
Later among the works it cites.
Motionctrl: A unified and flexible motion controller for video generation
Z. Wang, Z. Yuan, X. Wang, Y. Li, T. Chen, M. Xia, P. Luo, and Y. Shan · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence · 2023
Cited alongside, same era.
Instruct2act: Mapping multi-modality instructions to robotic actions with large language model
S. Huang, Z. Jiang, H. Dong, Y. Qiao, P. Gao, and H. Li · 2023
Cited alongside, same era.
Libero: Benchmarking knowledge transfer for lifelong robot learning
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Cited alongside, same era.
Daydreamer: World models for physical robot learning
P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg · 2023
Cited alongside, same era.
Genesis: A universal and generative physics engine for robotics and beyond, December 2024
G. Authors · 2024
Cited alongside, same era.
A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Cited alongside, same era.
Genie: Generative interactive environments
J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y. Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. d. Freitas, S. Singh, and T. Rocktäschel · 2024
Cited alongside, same era.
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al · 2024
Later among the works it cites.
Open-sora: Democratizing efficient video production for all, March 2024
Z. Zheng, X. Peng, T. Yang, C. Shen, S. Li, H. Liu, Y. Zhou, T. Li, and Y. You · 2024
Later among the works it cites.
Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras
A. Abouelenin, A. Ashfaq, A. Atkinson, H. Awadalla, N. Bach, J. Bao, A. Benhaim, M. Cai, V. Chaudhary, C. Chen, et al · 2025
Closest in time.
Cosmos world foundation model platform for physical AI
N. Agarwal, A. Ali, M. Bala, Y. Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y. Chen, Y. Cui, Y. Ding, D. Dworakowski, J. Fan, M. Fenzi, F. Ferroni, S. Fidler, D. Fox, S. Ge, Y. Ge, J. Gu, S. Gururani, E. He, J. Huang, J. Huffman, P. Jannaty, J. Jin, S. W. Kim, G. Klár, G. Lam, S. Lan, L. Leal-Taixe, A. Li, Z. Li, C.-H. Lin, T.-Y. Lin, H. Ling, M.-Y. Liu, X. Liu, A. Luo, Q. Ma, H. Mao, K. Mo, A. Mousavian, S. Nah, S. Niverty, D. Page, D. Paschalidou, Z. Patel, L. Pavao, M. Ramezanali, F. Reda, X. Ren, V. R. N. Sabavat, E. Schmerling, S. Shi, B. Stefaniak, S. Tang, L. Tchapmi, P. Tredak, W.-C. Tseng, J. Varghese, H. Wang, H. Wang, H. Wang, T.-C. Wang, F. Wei, X. Wei, J. Z. Wu, J. Xu, W. Yang, L. Yen-Chen, X. Zeng, Y. Zeng, J. Zhang, Q. Zhang, Y. Zhang, Q. Zhao, and A. Zolkowski · 2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
T. Chen, Z. Chen, B. Chen, Z. Cai, Y. Liu, Q. Liang, Z. Li, X. Lin, Y. Ge, Z. Gu, et al · 2025
Closest in time.
Enerverse: Envisioning embodied future space for robotics manipulation
S. Huang, L. Chen, P. Zhou, S. Chen, Z. Jiang, Y. Hu, Y. Liao, P. Gao, H. Li, M. Yao, et al · 2025
Closest in time.
Dreamgen: Unlocking generalization in robot learning through video world models
J. Jang, S. Ye, Z. Lin, J. Xiang, J. Bjorck, Y. Fang, F. Hu, S. Huang, K. Kundalia, Y.-C. Lin, et al · 2025
Closest in time.
Enerverse-ac: Envisioning embodied environments with action condition
Y. Jiang, S. Chen, S. Huang, L. Chen, P. Zhou, Y. Liao, X. He, C. Liu, H. Li, M. Yao, et al · 2025
Closest in time.
Gaia-2: A controllable multi-view generative world model for autonomous driving
L. Russell, A. Hu, L. Bertoni, G. Fedoseev, J. Shotton, E. Arani, and G. Corrado · 2025
Closest in time.
Ewmbench: Evaluating scene, motion, and semantic quality in embodied world models
H. Yue, S. Huang, Y. Liao, S. Chen, P. Zhou, L. Chen, M. Yao, and G. Ren · 2025
Closest in time.
Autoeval: Autonomous evaluation of generalist robot manipulation policies in the real world
Z. Zhou, P. Atreya, Y. L. Tan, K. Pertsch, and S. Levine · 2025
Closest in time.