Fetching the paper…
Reading the bibliography…
End-to-end imitation learning frameworks (e.g., VLA) are increasingly prominent in robotics, as they enable rapid task transfer by learning directly from perception to control, eliminating the need for complex hand-crafted features.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2010
Earlier work this paper cites.
Octnet: Learning deep 3d representations at high resolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3577–3586
Gernot Riegler, Ali Osman Ulusoy, and Andreas Geiger. 2017 · 2017
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2020 · 2020
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al · 2022
Earlier work this paper cites.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13739–13748
Stephen James, Kentaro Wada, Tristan Laidlow, and Andrew J Davison. 2022 · 2022
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning . PMLR, 991–1002
Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn. 2022 · 2022
Earlier work this paper cites.
Planning with diffusion for flexible behavior synthesis
Michael Janner, Yilun Du, Joshua B Tenenbaum, and Sergey Levine. 2022 · 2022
Earlier work this paper cites.
Polarnet: 3d point clouds for language-guided robotic manipulation
Shizhe Chen, Ricardo Garcia, Cordelia Schmid, and Ivan Laptev. 2023 · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. 2023 · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, et al · 2023
Earlier work this paper cites.
Foundation models in robotics: Applications, challenges, and the future
Roya Firoozi, Johnathan Tucker, Stephen Tian, Anirudha Majumdar, Jiankai Sun, Weiyu Liu, Yuke Zhu, Shuran Song, Ashish Kapoor, Karol Hausman, et al · 2023
Earlier work this paper cites.
One act play: Single demonstration behavior cloning with action chunking transformers
Abraham George and Amir Barati Farimani. 2023 · 2023
Earlier work this paper cites.
Act3d: 3d feature field transformers for multi-task robotic manipulation
Theophile Gervet, Zhou Xian, Nikolaos Gkanatsios, and Katerina Fragkiadaki. 2023 · 2023
Earlier work this paper cites.
Rvt: Robotic view transformer for 3d object manipulation. In Conference on Robot Learning . PMLR, 694–710
Ankit Goyal, Jie Xu, Yijie Guo, Valts Blukis, Yu-Wei Chao, and Dieter Fox. 2023 · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. 2023 · 2023
Earlier work this paper cites.
Instruction-driven history-aware policies for robotic manipulations. In Conference on Robot Learning . PMLR, 175–187
Pierre-Louis Guhur, Shizhe Chen, Ricardo Garcia Pinel, Makarand Tapaswi, Ivan Laptev, and Cordelia Schmid. 2023 · 2023
Earlier work this paper cites.
Scaling up and distilling down: Language-guided robot skill acquisition. In Conference on Robot Learning . PMLR, 3766–3777
Huy Ha, Pete Florence, and Shuran Song. 2023 · 2023
Earlier work this paper cites.
Papras: Plug-and-play robotic arm system
Joohyung Kim, Dhruv C Mathur, Kazuki Shin, and Sean Taylor. 2023 · 2023
Earlier work this paper cites.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023 · 2023
Earlier work this paper cites.
Perceiver-actor: A multi-task transformer for robotic manipulation. In Conference on Robot Learning . PMLR, 785–799
Mohit Shridhar, Lucas Manuelli, and Dieter Fox. 2023 · 2023
Earlier work this paper cites.
Learning fine-grained bimanual manipulation with low-cost hardware
Tony Z Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. 2023 · 2023
Earlier work this paper cites.
π 0 \pi_{0} : A Vision-Language-Action Flow Model for General Robot Control
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky. 2024 · 2024
Earlier work this paper cites.
Motion planning diffusion: Learning and adapting robot motion planning with diffusion models
Joao Carvalho, A Le, Piotr Kicki, Dorothea Koert, and Jan Peters. 2024 · 2024
Earlier work this paper cites.
Meshxl: Neural coordinate field for generative 3d foundation models
Sijin Chen, Xin Chen, Anqi Pang, Xianfang Zeng, Wei Cheng, Yijun Fu, Fukun Yin, Billzb Wang, Jingyi Yu, Gang Yu, et al · 2024
Cited alongside, same era.
Demonstrating a Robust Walking Algorithm for Underactuated Bipedal Robots in Non-flat, Non-stationary Environments. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 11210–11217
Oluwami Dosunmu-Ogunbi, Aayushi Shrivastava, and Jessy W Grizzle. 2024 · 2024
Cited alongside, same era.
Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 653–660
Hao-Shu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, and Cewu Lu. 2024 · 2024
Cited alongside, same era.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Zipeng Fu, Tony Z Zhao, and Chelsea Finn. 2024 · 2024
Cited alongside, same era.
Humanvla: Towards vision-language directed object rearrangement by physical humanoid
Xinyu Xu, Yizheng Zhang, Yong-Lu Li, Lei Han, and Cewu Lu. 2024 · 2024
Later among the works it cites.
Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution
Yang Yue, Yulin Wang, Bingyi Kang, Yizeng Han, Shenzhi Wang, Shiji Song, Jiashi Feng, and Gao Huang. 2024 · 2024
Later among the works it cites.
Generalizable humanoid manipulation with improved 3d diffusion policies
Yanjie Ze, Zixuan Chen, Wenhao Wang, Tianyi Chen, Xialin He, Ying Yuan, Xue Bin Peng, and Jiajun Wu. 2024 · 2024
Later among the works it cites.
Sparsevlm: Visual token sparsification for efficient vision-language model inference
Yuan Zhang, Chun-Kai Fan, Junpeng Ma, Wenzhao Zheng, Tao Huang, Kuan Cheng, Denis Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Copa: General robotic manipulation through spatial constraints of parts with foundation models. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 9488–9495
Haoxu Huang, Fanqi Lin, Yingdong Hu, Shengjie Wang, and Yang Gao. 2024 · 2024
Cited alongside, same era.
Robust Imitation Learning for Mobile Manipulator Focusing on Task-Related Viewpoints and Regions. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2885–2892
Yutaro Ishida, Yuki Noguchi, Takayuki Kanai, Kazuhiro Shintani, and Hiroshi Bito. 2024 · 2024
Cited alongside, same era.
A Survey on Vision Autoregressive Model
Kai Jiang and Jiaxing Huang. 2024 · 2024
Cited alongside, same era.
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al · 2024
Cited alongside, same era.
Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo, Xi Chen, Mozheng Liao, Fangyun Wei, Yu Deng, Sicheng Xu, Yizhong Zhang, et al · 2024
Cited alongside, same era.
Evaluating Real-World Robot Manipulation Policies in Simulation
Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat, Isabel Sieh, Sean Kirmani, Sergey Levine, Jiajun Wu, Chelsea Finn, Hao Su, Quan Vuong, and Ted Xiao. 2024a · 2024
Cited alongside, same era.
Llara: Supercharging robot learning data for vision-language policy
Xiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan Burgert, Mu Cai, Yong Jae Lee, et al · 2024
Cited alongside, same era.
Experience-Learning Inspired Two-Step Reward Method for Efficient Legged Locomotion Learning Towards Natural and Robust Gaits. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 13297–13302
Yinghui Li, Jinze Wu, Xin Liu, Weizhong Guo, and Yufei Xue. 2024d · 2024
Cited alongside, same era.
Wangbo Zhao, Jiasheng Tang, Yizeng Han, Yibing Song, Kai Wang, Gao Huang, Fan Wang, and Yang You. 2024 · 2024
Later among the works it cites.
3d-vla: A 3d vision-language-action generative world model
Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan. 2024 · 2024
Later among the works it cites.
Survey on Vision-Language-Action Models
Adilzhan Adilkhanov, Amir Yelenov, Assylkhan Seitzhanov, Ayan Mazhitov, Azamat Abdikarimov, Danissa Sandykbayeva, Daryn Kenzhebek, Dinmukhammed Mukashev, Ilyas Umurbekov, Jabrail Chumakov, et al · 2025
Closest in time.
Mixture of Action Expert Embeddings: Multi-Task ACT
Suhyung Choi, Youngseok Joo, Jun Ki Lee, and Byoung-Tak Zhang. 2025 · 2025
Closest in time.
Redefining Robot Generalization Through Interactive Intelligence
Sharmita Dey. 2025 · 2025
Closest in time.
AgentRefine: Enhancing Agent Generalization through Refinement Tuning
Dayuan Fu, Keqing He, Yejie Wang, Wentao Hong, Zhuoma Gongque, Weihao Zeng, Wei Wang, Jingang Wang, Xunliang Cai, and Weiran Xu. 2025 · 2025
Closest in time.
Improving Vision-Language-Action Model with Online Reinforcement Learning
Yanjiang Guo, Jianke Zhang, Xiaoyu Chen, Xiang Ji, Yen-Jen Wang, Yucheng Hu, and Jianyu Chen. 2025 · 2025
Closest in time.
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
Zhi Hou, Tianyi Zhang, Yuwen Xiong, Haonan Duan, Hengjun Pu, Ronglei Tong, Chengyang Zhao, Xizhou Zhu, Yu Qiao, Jifeng Dai, et al · 2025
Closest in time.
JAKA Robots
JAKA. 2025 · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success
Moo Jin Kim, Chelsea Finn, and Percy Liang. 2025 · 2025
Closest in time.
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Guangxuan Xiao, and Song Han. 2025 · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. 2025 · 2025
Closest in time.
Trossen Robotics
Trossen Robotics. 2025 · 2025
Closest in time.
Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
Wenxuan Song, Jiayi Chen, Pengxiang Ding, Han Zhao, Wei Zhao, Zhide Zhong, Zongyuan Ge, Jun Ma, and Haoang Li. 2025 · 2025
Closest in time.
Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu, Zhibin Tang, Kun Wu, Zhiyuan Xu, Ning Liu, Ran Cheng, Chaomin Shen, et al · 2025
Closest in time.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al · 2025
Closest in time.
Siyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu, Tao Huang, and Chang Xu. 2025 · 2025
Closest in time.
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Jianke Zhang, Yanjiang Guo, Yucheng Hu, Xiaoyu Chen, Xiang Zhu, and Jianyu Chen. 2025a · 2025
Closest in time.
Autoregressive action sequence learning for robotic manipulation
Xinyu Zhang, Yuhan Liu, Haonan Chang, Liam Schramm, and Abdeslam Boularias. 2025b · 2025
Closest in time.
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, et al · 2025
Closest in time.