Fetching the paper…
Reading the bibliography…
Video Generation Models (VGMs) have become powerful backbones for Vision-Language-Action (VLA) models, leveraging large-scale pretraining for robust dynamics modeling.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2017 · 2017
Earlier work this paper cites.
Ha, D.; and Schmidhuber, J. 2018 · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
FVD: A new metric for video generation
Unterthiner, T.; Van Steenkiste, S.; Kurach, K.; Marinier, R.; Michalski, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
James, S.; Ma, Z.; Arrojo, D. R.; and Davison, A. J. 2020 · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning
Ramos, S.; Girgin, S.; Hussenot, L.; Vincent, D.; Yakubovich, H.; Toyama, D.; Gergely, A.; Stanczyk, P.; Marinier, R.; Harmsen, J.; Pietquin, O.; and Momchev, N. 2021 · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J.; and Salimans, T. 2022 · 2022
Earlier work this paper cites.
Scalable Diffusion Models with Transformers
Peebles, W.; and Xie, S. 2022 · 2022
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; et al. 2023 · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; and Song, S. 2023 · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D.; Xia, F.; Sajjadi, M. S.; Lynch, C.; Chowdhery, A.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; Huang, W.; et al. 2023 · 2023
Cited alongside, same era.
Rvt: Robotic view transformer for 3d object manipulation
Goyal, A.; Xu, J.; Guo, Y.; Blukis, V.; Chao, Y.-W.; and Fox, D. 2023 · 2023
Cited alongside, same era.
Daydreamer: World models for physical robot learning
Wu, P.; Escontrela, A.; Hafner, D.; Abbeel, P.; and Goldberg, K. 2023 · 2023
Cited alongside, same era.
Dynamicrafter: Animating open-domain images with video diffusion priors
Xing, J.; Xia, M.; Zhang, Y.; Chen, H.; Wang, X.; Wong, T.-T.; and Shan, Y. 2023 · 2023
Cited alongside, same era.
Video generation models as world simulators
OpenAI. 2024 · 2024
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
O’Neill, A.; Rehman, A.; Maddukuri, A.; Gupta, A.; Padalkar, A.; Lee, A.; Pooley, A.; Gupta, A.; Mandlekar, A.; Jain, A.; et al. 2024 · 2024
Later among the works it cites.
RoboDreamer: Learning compositional world models for robot imagination
Zhou, S.; Du, Y.; Chen, J.; Li, Y.; Yeung, D.-Y.; and Gan, C. 2024 · 2024
Later among the works it cites.
Cosmos world foundation model platform for physical ai
Agarwal, N.; Ali, A.; Bala, M.; Balaji, Y.; Barker, E.; Cai, T.; Chattopadhyay, P.; Chen, Y.; Cui, Y.; Ding, Y.; et al. 2025 · 2025
Closest in time.
Univla: Learning to act anywhere with task-centric latent actions
Bu, Q.; Yang, Y.; Cai, J.; Gao, S.; Ren, G.; Yao, M.; Luo, P.; and Li, H. 2025 · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang, M.; Du, Y.; Ghasemipour, K.; Tompson, J.; Schuurmans, D.; and Abbeel, P. 2023 · 2023
Cited alongside, same era.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
Cheang, C.-L.; Chen, G.; Jing, Y.; Kong, T.; Li, H.; Li, Y.; Liu, Y.; Wu, H.; Xu, J.; Yang, Y.; et al. 2024 · 2024
Cited alongside, same era.
EVA: An Embodied World Model for Future Video Anticipation
Chi, X.; Zhang, H.; Fan, C.-K.; Qi, X.; Zhang, R.; Chen, A.; Chan, C.-m.; Xue, W.; Luo, W.; Zhang, S.; et al. 2024 · 2024
Cited alongside, same era.
OpenVLA: An Open-Source Vision-Language-Action Model
Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; Vuong, Q.; Kollar, T.; Burchfiel, B.; Tedrake, R.; Sadigh, D.; Levine, S.; Liang, P.; and Finn, C. 2024 · 2024
Cited alongside, same era.
Li, Q.; Liang, Y.; Wang, Z.; Luo, L.; Chen, X.; Liao, M.; Wei, F.; Deng, Y.; Xu, S.; Zhang, Y.; Wang, X.; Liu, B.; Fu, J.; Bao, J.; Chen, D.; Shi, Y.; Yang, J.; and Guo, B. 2024 · 2024
Cited alongside, same era.
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
Liu, J.; Liu, M.; Wang, Z.; An, P.; Li, X.; Zhou, K.; Yang, S.; Zhang, R.; Guo, Y.; and Zhang, S. 2024 · 2024
Cited alongside, same era.
π 0 \pi_{0} : A Vision-Language-Action Flow Model for General Robot Control
Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; Jakubczak, S.; Jones, T.; Ke, L.; Levine, S.; Li-Bell, A.; Mothukuri, M.; Nair, S.; Pertsch, K.; Shi, L. X.; Tanner, J.; Vuong, Q.; Walling, A.; Wang, H.; and Zhilinsky, U. 2024a
Cited in the paper.
Closest in time.
Adaworld: Learning adaptable world models with latent actions
Gao, S.; Zhou, S.; Du, Y.; Zhang, J.; and Gan, C. 2025 · 2025
Closest in time.
Enerverse: Envisioning embodied future space for robotics manipulation
Huang, S.; Chen, L.; Zhou, P.; Chen, S.; Jiang, Z.; Hu, Y.; Liao, Y.; Gao, P.; Li, H.; Yao, M.; et al. 2025 · 2025
Closest in time.
Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model
Liu, J.; Chen, H.; An, P.; Liu, Z.; Zhang, R.; Gu, C.; Li, X.; Guo, Z.; Chen, S.; Liu, M.; et al. 2025 · 2025
Closest in time.
Unified Vision-Language-Action Model
Wang, Y.; Li, X.; Wang, W.; Zhang, J.; Li, Y.; Chen, Y.; Wang, X.; and Zhang, Z. 2025 · 2025
Closest in time.
Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
Wen, J.; Zhu, Y.; Li, J.; Zhu, M.; Tang, Z.; Wu, K.; Xu, Z.; Liu, N.; Cheng, R.; Shen, C.; et al. 2025 · 2025
Closest in time.
Zhang, R.; Dong, M.; Zhang, Y.; Heng, L.; Chi, X.; Dai, G.; Du, L.; Du, Y.; and Zhang, S. 2025 · 2025
Closest in time.