Fetching the paper…
Reading the bibliography…
Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which, despite strong generalization in static manipulation, struggle in dynamic scenarios requiring rapid perception, temporal anticipation, and continuous control.
RoboCup: the robot world cup initiative
Hiroaki Kitano, Minoru Asada, Yasuo Kuniyoshi, Itsuki Noda, and Eiichi Osawa · 1997
Earlier work this paper cites.
ROBOTURK: A crowdsourcing platform for robotic skill learning through imitation
Ajay Mandlekar, Yuke Zhu, Animesh Garg, Jonathan Booher, Max Spero, Albert Tung, Julian Gao, John Emmons, Anchit Gupta, Emre Orbay, Silvio Savarese, and Li Fei-Fei · 2018
Earlier work this paper cites.
TossingBot: learning to throw arbitrary objects with residual physics
Andy Zeng, Shuran Song, Johnny Lee, Alberto Rodriguez, and Thomas A. Funkhouser · 2020
Earlier work this paper cites.
3D-FRONT: 3D furnished rooms with layouts and semantics
Huan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang, Jiaming Wang, Cao Li, Qixun Zeng, Chengyue Sun, Rongfei Jia, Binqiang Zhao, and Hao Zhang · 2021
Earlier work this paper cites.
COYO-700M: image-text pair dataset
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim · 2022
Earlier work this paper cites.
VIMA: general robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan · 2022
Earlier work this paper cites.
CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Oier Mees, Lukás Hermann, Erick Rosete-Beas, and Wolfram Burgard · 2022
Earlier work this paper cites.
Efficient tactile simulation with differentiability for robotic manipulation
Jie Xu, Sangwoon Kim, Tao Chen, Alberto Rodriguez Garcia, Pulkit Agrawal, Wojciech Matusik, and Shinjiro Sueda · 2022
Earlier work this paper cites.
VLMbench: A compositional benchmark for vision-and-language manipulation
Kaizhi Zheng, Xiaotong Chen, Odest Chadwicke Jenkins, and Xin Eric Wang · 2022
Earlier work this paper cites.
RT-1: robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alexander Herzog, and Jasmine Hsu et al · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song · 2023
Earlier work this paper cites.
Objaverse: A universe of annotated 3D objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi · 2023
Earlier work this paper cites.
Flow matching for generative modeling
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le · 2023
Earlier work this paper cites.
GPT-4 technical report
OpenAI · 2023
Earlier work this paper cites.
FastViT: A fast hybrid vision transformer using structural reparameterization
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel, and Anurag Ranjan · 2023
Earlier work this paper cites.
BridgeData V2: A dataset for robot learning at scale
Homer Rich Walke, Kevin Black, Tony Z. Zhao, Quan Vuong, Chongyi Zheng, Philippe Hansen-Estruch, Andre Wang He, Vivek Myers, Moo Jin Kim, Max Du, Abraham Lee, Kuan Fang, Chelsea Finn, and Sergey Levine · 2023
Earlier work this paper cites.
Learning fine-grained bimanual manipulation with low-cost hardware
Tony Z. Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn · 2023
Earlier work this paper cites.
RT-2: Vision-Language-Action models transfer web knowledge to robotic control
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, and Ayzaan Wahid et al · 2023
Earlier work this paper cites.
Achieving human level competitive robot table tennis
David B. D’Ambrosio, Saminda Abeyruwan, Laura Graesser, Atil Iscen, Heni Ben Amor, Alex Bewley, Barney J. Reed, Krista Reymann, Leila Takayama, and Yuval Tassa et al · 2024
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach · 2024
Cited alongside, same era.
Octo: An open-source generalist robot policy
Dibya Ghosh, Homer Rich Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yunliang Chen, Quan Vuong, Ted Xiao, Pannag R. Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine · 2024
Cited alongside, same era.
OpenVLA: An open-source Vision-Language-Action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Paul Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn · 2024
Cited alongside, same era.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2024
π \pi 0 {}_{\mbox{0}} : A Vision-Language-Action flow model for general robot control
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, and Brian Ichter et al · 2025
Later among the works it cites.
π \pi 0.5 {}_{\mbox{0.5}} : a Vision-Language-Action model with open-world generalization
Physical Intelligence · 2025
Later among the works it cites.
Fine-tuning Vision-Language-Action models: Optimizing speed and success
Moo Jin Kim, Chelsea Finn, and Percy Liang · 2025
Later among the works it cites.
Train once, deploy anywhere: Realize data-efficient dynamic object manipulation
Zhuoling Li, Xiaoyang Wu, Zhenhua Xu, and Hengshuang Zhao · 2025
Later among the works it cites.
Running VLAs at real-time speed
Yunchao Ma, Yizhuang Zhou, Yunhuan Yang, Tiancai Wang, and Haoqiang Fan · 2025
Later among the works it cites.
SmolVLM: redefining small and efficient multimodal models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
RoboCasa: Large-scale simulation of everyday tasks for generalist robots
Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Mandlekar, and Yuke Zhu · 2024
Cited alongside, same era.
Open X-Embodiment: robotic learning datasets and RT-X models
Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, and Ajinkya Jain et al · 2024
Cited alongside, same era.
The llama 3 herd of models
Llama Team · 2024
Cited alongside, same era.
RoboGen: Towards unleashing infinite data for automated robot learning via generative simulation
Yufei Wang, Zhou Xian, Feng Chen, Tsun-Hsuan Wang, Yian Wang, Katerina Fragkiadaki, Zackory Erickson, David Held, and Chuang Gan · 2024
Cited alongside, same era.
Unleashing large-scale video generative pre-training for visual robot manipulation
Hongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Hang Li, and Tao Kong · 2024
Cited alongside, same era.
Efficient track anything
Yunyang Xiong, Chong Zhou, Xiaoyu Xiang, Lemeng Wu, Chenchen Zhu, Zechun Liu, Saksham Suri, Balakrishnan Varadarajan, Ramya Akula, Forrest N. Iandola, Raghuraman Krishnamoorthi, Bilge Soran, and Vikas Chandra · 2024
Cited alongside, same era.
Learning interactive real-world simulators
Sherry Yang, Yilun Du, Seyed Kamyar Seyed Ghasemipour, Jonathan Tompson, Leslie Pack Kaelbling, Dale Schuurmans, and Pieter Abbeel · 2024
Cited alongside, same era.
Andrés Marafioti, Orr Zohar, Miquel Farré, Merve Noyan, Elie Bakouch, Pedro Cuenca, Cyril Zakka, Loubna Ben Allal, Anton Lozhkov, Nouamane Tazi, Vaibhav Srivastav, Joshua Lochner, Hugo Larcher, Mathieu Morlon, Lewis Tunstall, Leandro von Werra, and Thomas Wolf · 2025
Later among the works it cites.
Isaac Lab: A gpu-accelerated simulation framework for multi-modal robot learning
Mayank Mittal, Pascal Roth, James Tigue, Antoine Richard, Octi Zhang, Peter Du, Antonio Serrano-Muñoz, Xinjie Yao, René Zurbrügg, and Nikita Rudin et al · 2025
Later among the works it cites.
Fine-tuning Vision-Language-Action models: Optimizing speed and success
Percy Liang Moo Jin Kim, Chelsea Finn · 2025
Later among the works it cites.
RoboTwin: dual-arm robot benchmark with generative digital twins
Yao Mu, Tianxing Chen, Zanxin Chen, Shijia Peng, Zhiqian Lan, Zeyu Gao, Zhixuan Liang, Qiaojun Yu, Yude Zou, Mingkun Xu, Lunkai Lin, Zhiqiang Xie, Mingyu Ding, and Ping Luo · 2025
Later among the works it cites.
VideoWorld: Exploring knowledge learning from unlabeled videos
Zhongwei Ren, Yunchao Wei, Xun Guo, Yao Zhao, Bingyi Kang, Jiashi Feng, and Xiaojie Jin · 2025
Later among the works it cites.
SmolVLA: A Vision-Language-Action model for affordable and efficient robotics
Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andrés Marafioti, Simon Alibert, Matthieu Cord, Thomas Wolf, and Rémi Cadène · 2025
Later among the works it cites.
VLASH: real-time VLAs via future-state-aware asynchronous inference
Jiaming Tang, Yufei Sun, Yilong Zhao, Shang Yang, Yujun Lin, Zhuoyang Zhang, James Hou, Yao Lu, Zhijian Liu, and Song Han · 2025
Later among the works it cites.
VLA-Adapter: an effective paradigm for tiny-scale Vision-Language-Action model
Yihao Wang, Pengxiang Ding, Lingxiao Li, Can Cui, Zirui Ge, Xinyang Tong, Wenxuan Song, Han Zhao, Wei Zhao, Pengxu Hou, Siteng Huang, Yifan Tang, Wenhui Wang, Ru Zhang, Jianyi Liu, and Donglin Wang · 2025
Later among the works it cites.
3D scene generation: A survey
Beichen Wen, Haozhe Xie, Zhaoxi Chen, Fangzhou Hong, and Ziwei Liu · 2025
Later among the works it cites.
RoboMIND: benchmark on multi-embodiment intelligence normative data for robot manipulation
Kun Wu, Chengkai Hou, Jiaming Liu, Zhengping Che, Xiaozhu Ju, Zhuqin Yang, Meng Li, Yinuo Zhao, Zhiyuan Xu, and Guang Yang et al · 2025
Later among the works it cites.
Dynamic behavior cloning with temporal feature prediction: Enhancing robotic arm manipulation in moving object tasks
Yifan Zhang, Ruiping Wang, and Xilin Chen · 2025
Later among the works it cites.
A survey on Vision-Language-Action models: An action tokenization perspective
Yifan Zhong, Fengshuo Bai, Shaofei Cai, Xuchuan Huang, Zhang Chen, Xiaowei Zhang, Yuanfei Wang, Shaoyang Guo, Tianrui Guan, Ka Nam Lui, Zhiquan Qi, Yitao Liang, Yuanpei Chen, and Yaodong Yang · 2025
Later among the works it cites.