Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics.
Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, et al · 2018
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, et al · 2022
Earlier work this paper cites.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wolfram Burgard · 2022
Earlier work this paper cites.
Playfusion: Skill acquisition via diffusion from language-annotated play
Lili Chen, Shikhar Bahl, and Deepak Pathak · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, et al · 2023
Earlier work this paper cites.
Maniskill2: A unified benchmark for generalizable manipulation skills
Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, et al · 2023
Earlier work this paper cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Earlier work this paper cites.
LIBERO: Benchmarking knowledge transfer for lifelong robot learning
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone · 2023
Earlier work this paper cites.
Latent plans for task-agnostic offline reinforcement learning
Erick Rosete-Beas, Oier Mees, Gabriel Kalweit, Joschka Boedecker, and Wolfram Burgard · 2023
Earlier work this paper cites.
Mutex: Learning unified policies from multimodal task specifications
Rutav Shah, Roberto Martín-Martín, and Yuke Zhu · 2023
Earlier work this paper cites.
Bridgedata v2: A dataset for robot learning at scale
Homer Rich Walke, Kevin Black, Tony Z Zhao, et al · 2023
Earlier work this paper cites.
Train offline, test online: A real robot learning benchmark
Gaoyue Zhou, Victoria Dean, Mohan Kumar Srirama, Aravind Rajeswaran, Jyothish Pari, Kyle Hatch, Aryan Jain, Tianhe Yu, Pieter Abbeel, Lerrel Pinto, et al · 2023
Earlier work this paper cites.
Viola: Imitation learning for vision-based manipulation with object proposal priors
Yifeng Zhu, Abhishek Joshi, Peter Stone, and Yuke Zhu · 2023
Earlier work this paper cites.
RT-2: Vision-language-action models transfer web knowledge to robotic control
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, et al · 2023
Earlier work this paper cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
Kevin Black, Noah Brown, Danny Driess, et al · 2024
Earlier work this paper cites.
Genie: Generative interactive environments
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, et al · 2024
Cited alongside, same era.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, et al · 2024
Cited alongside, same era.
Berkeley UR5 demonstration dataset, 2024
Lawrence Yunliang Chen, Simeon Adebola, and Ken Goldberg · 2024
Cited alongside, same era.
OpenVLA: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, et al · 2024
Cited alongside, same era.
Open-sora plan: Open-source large video generation model
Bin Lin, Yunyang Ge, Xinhua Cheng, et al · 2024
Cited alongside, same era.
Dita: Scaling diffusion transformer for generalist vision-language-action policy
Zhi Hou, Tianyi Zhang, Yuwen Xiong, Haonan Duan, et al · 2025
Later among the works it cites.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization
Physical Intelligence, Kevin Black, Noah Brown, et al · 2025
Later among the works it cites.
Fine-tuning vision-language-action models: Optimizing speed and success
Moo Jin Kim, Chelsea Finn, and Percy Liang · 2025
Later among the works it cites.
Disentangled motion modeling for video frame interpolation
Jaihyun Lew, Jooyoung Choi, Chaehun Shin, Dahuin Jung, and Sungroh Yoon · 2025
Later among the works it cites.
Unified video action model
Shuang Li, Yihuai Gao, Dorsa Sadigh, and Shuran Song · 2025
Later among the works it cites.
Fast: Efficient action tokenization for vision-language-action models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration
Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, et al · 2024
Cited alongside, same era.
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling
Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, et al · 2024
Cited alongside, same era.
Eggen: Image generation with multi-entity prior learning through entity guidance
Zhenhong Sun, Junyan Wang, Zhiyu Tan, Daoyi Dong, Hailan Ma, Hao Li, and Dong Gong · 2024
Cited alongside, same era.
Occllama: An occupancy-language-action generative world model for autonomous driving
Julong Wei, Shanshuai Yuan, Pengfei Li, Qingda Hu, Zhongxue Gan, and Wenchao Ding · 2024
Cited alongside, same era.
Efficient video diffusion models via content-frame motion-latent decomposition
Sihyun Yu, Weili Nie, De-An Huang, Boyi Li, Jinwoo Shin, and Anima Anandkumar · 2024
Cited alongside, same era.
V-jepa 2: Self-supervised video models enable understanding, prediction and planning
Mido Assran, Adrien Bardes, David Fan, et al · 2025
Cited alongside, same era.
Gr00t n1: An open foundation model for generalist humanoid robots
Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, et al · 2025
Cited alongside, same era.
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, et al · 2025
Later among the works it cites.
Spatialvla: Exploring spatial representations for visual-language-action model
Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, et al · 2025
Later among the works it cites.
Llapa: A vision-language model framework for counterfactual-aware procedural planning
Shibo Sun, Xue Li, Donglin Di, Mingjie Wei, Lanshun Nie, Wei-Nan Zhang, Dechen Zhan, Yang Song, and Lei Fan · 2025
Later among the works it cites.
Wan: Open and advanced large-scale video generative models
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, et al · 2025
Later among the works it cites.
Vidtwin: Video vae with decoupled structure and dynamics
Yuchi Wang, Junliang Guo, Xinyi Xie, Tianyu He, Xu Sun, and Jiang Bian · 2025
Later among the works it cites.
Latent action pretraining from videos
Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Sejune Joo, Jianwei Yang, et al · 2025
Later among the works it cites.
Dreamvla: A vision-language-action model dreamed with comprehensive world knowledge
Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang, et al · 2025
Later among the works it cites.
Cot-VLA: Visual chain-of-thought reasoning for vision-language-action models
Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, et al · 2025
Later among the works it cites.
Flowvla: Visual chain of thought-based motion reasoning for vision-language-action models
Zhide Zhong, Haodong Yan, Junfeng Li, Xiangchen Liu, Xin Gong, et al · 2025
Later among the works it cites.
Vipra: Video prediction for robot actions
Sandeep Routray, Hengkai Pan, Unnat Jain, Shikhar Bahl, and Deepak Pathak · 2026
Closest in time.