Fetching the paper…
Reading the bibliography…
Recent vision-language-action (VLA) models built on pretrained vision-language models (VLMs) have demonstrated strong performance in robotic manipulation.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V.; Kulkarni, G.; and Berg, T. 2011 · 2011
Earlier work this paper cites.
Long short-term memory
Graves, A.; and Graves, A. 2012 · 2012
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S.; Ordonez, V.; Matten, M.; and Berg, T. 2014 · 2014
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D.; Irpan, A.; Pastor, P.; Ibarz, J.; Herzog, A.; Jang, E.; Quillen, D.; Holly, E.; Kalakrishnan, M.; Vanhoucke, V.; et al. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019 · 2019
Earlier work this paper cites.
Rlds: an ecosystem to generate, share and use datasets in reinforcement learning
Ramos, S.; Girgin, S.; Hussenot, L.; Vincent, D.; Yakubovich, H.; Toyama, D.; Gergely, A.; Stanczyk, P.; Marinier, R.; Harmsen, J.; et al. 2021 · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Dabis, J.; Finn, C.; Gopalakrishnan, K.; Hausman, K.; Herzog, A.; Hsu, J.; et al. 2022 · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. 2022 · 2022
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A.; Brown, N.; Carbajal, J.; Chebotar, Y.; Chen, X.; Choromanski, K.; Ding, T.; Driess, D.; Dubey, A.; Finn, C.; et al. 2023 · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; and Song, S. 2023 · 2023
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
Driess, D.; Xia, F.; Sajjadi, M. S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. 2023 · 2023
Earlier work this paper cites.
Vision-Language Foundation Models as Effective Robot Imitators
Li, X.; Liu, M.; Zhang, H.; Yu, C.; Xu, J.; Wu, H.; Cheang, C.; Jing, Y.; Zhang, W.; Liu, H.; Li, H.; and Kong, T. 2023 · 2023
Earlier work this paper cites.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Cited alongside, same era.
Interactive language: Talking to robots in real time
Lynch, C.; Wahid, A.; Tompson, J.; Ding, T.; Betker, J.; Baruch, R.; Armstrong, T.; and Florence, P. 2023 · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. 2023 · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W.; and Xie, S. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Cited alongside, same era.
OpenVLA: An Open-Source Vision-Language-Action Model
Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; et al. 2024 · 2024
Later among the works it cites.
Libero: Benchmarking knowledge transfer for lifelong robot learning
Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; and Stone, P. 2024 · 2024
Later among the works it cites.
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers
Ma, N.; Goldstein, M.; Albergo, M. S.; Boffi, N. M.; Vanden-Eijnden, E.; and Xie, S. 2024 · 2024
Later among the works it cites.
Octo: An Open-Source Generalist Robot Policy
Octo Model Team; Ghosh, D.; Walke, H.; Pertsch, K.; Black, K.; Mees, O.; Dasari, S.; Hejna, J.; Xu, C.; Luo, J.; Kreiman, T.; Tan, Y.; Sanketi, P.; Vuong, Q.; Xiao, T.; Sadigh, D.; Finn, C.; and Levine, S. 2024 · 2024
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
O’Neill, A.; Rehman, A.; Maddukuri, A.; Gupta, A.; Padalkar, A.; Lee, A.; Pooley, A.; Gupta, A.; Mandlekar, A.; Jain, A.; et al. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bridgedata v2: A dataset for robot learning at scale
Walke, H. R.; Black, K.; Zhao, T. Z.; Vuong, Q.; Zheng, C.; Hansen-Estruch, P.; He, A. W.; Myers, V.; Kim, M. J.; Du, M.; et al. 2023 · 2023
Cited alongside, same era.
Unleashing large-scale video generative pre-training for visual robot manipulation
Wu, H.; Jing, Y.; Cheang, C.; Chen, G.; Xu, J.; Li, X.; Liu, M.; Li, H.; and Kong, T. 2023 · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L. 2023 · 2023
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
Zhao, T. Z.; Kumar, V.; Levine, S.; and Finn, C. 2023 · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M.; Aneja, J.; Awadalla, H.; Awadallah, A.; Awan, A. A.; Bach, N.; Bahree, A.; Bakhtiari, A.; Bao, J.; Behl, H.; et al. 2024 · 2024
Cited alongside, same era.
MiniVLA: A Better VLA with a Smaller Footprint
Belkhale, S.; and Sadigh, D. 2024 · 2024
Cited alongside, same era.
π 0 \pi_{0} : A Vision-Language-Action Flow Model for General Robot Control
Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; Jakubczak, S.; Jones, T.; Ke, L.; Levine, S.; Li-Bell, A.; Mothukuri, M.; Nair, S.; Pertsch, K.; Shi, L. X.; Tanner, J.; Vuong, Q.; Walling, A.; Wang, H.; and Zhilinsky, U. 2024 · 2024
Cited alongside, same era.
Later among the works it cites.
Predictive inverse dynamics models are scalable learners for robotic manipulation
Tian, Y.; Yang, S.; Zeng, J.; Wang, P.; Lin, D.; Dong, H.; and Pang, J. 2024 · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; et al. 2024 · 2024
Later among the works it cites.
Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation
Wen, J.; Zhu, Y.; Li, J.; Zhu, M.; Wu, K.; Xu, Z.; Liu, N.; Cheng, R.; Shen, C.; Peng, Y.; et al. 2024 · 2024
Later among the works it cites.
3d diffusion policy
Ze, Y.; Zhang, G.; Zhang, K.; Hu, C.; Wang, M.; and Xu, H. 2024 · 2024
Later among the works it cites.
3d-vla: A 3d vision-language-action generative world model
Zhen, H.; Qiu, X.; Chen, P.; Yang, J.; Yan, X.; Du, Y.; Hong, Y.; and Gan, C. 2024 · 2024
Later among the works it cites.
Tracevla: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
Zheng, R.; Liang, Y.; Huang, S.; Gao, J.; Daumé III, H.; Kolobov, A.; Huang, F.; and Yang, J. 2024 · 2024
Later among the works it cites.
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Bjorck, J.; Castañeda, F.; Cherniadev, N.; Da, X.; Ding, R.; Fan, L.; Fang, Y.; Fox, D.; Hu, F.; Huang, S.; et al. 2025 · 2025
Closest in time.
Dita: Scaling diffusion transformer for generalist vision-language-action policy
Hou, Z.; Zhang, T.; Xiong, Y.; Duan, H.; Pu, H.; Tong, R.; Zhao, C.; Zhu, X.; Qiao, Y.; Dai, J.; et al. 2025 · 2025
Closest in time.
\ \backslash pi_ { \{ 0.5 } \} : a Vision-Language-Action Model with Open-World Generalization
Intelligence, P.; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; et al. 2025 · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success
Kim, M. J.; Finn, C.; and Liang, P. 2025 · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
Pertsch, K.; Stachowicz, K.; Ichter, B.; Driess, D.; Nair, S.; Vuong, Q.; Mees, O.; Finn, C.; and Levine, S. 2025 · 2025
Closest in time.
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Qu, D.; Song, H.; Chen, Q.; Yao, Y.; Ye, X.; Ding, Y.; Wang, Z.; Gu, J.; Zhao, B.; Wang, D.; et al. 2025 · 2025
Closest in time.