Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in visuomotor control, yet ensuring their robustness in unstructured real-world environments remains a persistent challenge.
Attention is all you need. advances in neural information processing systems
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Earlier work this paper cites.
Vima: General robot manipulation with multimodal prompts
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan · 2022
Earlier work this paper cites.
Data augmentation for manipulation, 2022
P. Mitrano and D. Berenson · 2022
Earlier work this paper cites.
Palm-e: An embodied multimodal language model, 2023
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Earlier work this paper cites.
Open x-embodiment: Robotic learning datasets and rt-x models
A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, et al · 2023
Earlier work this paper cites.
Bridgedata v2: A dataset for robot learning at scale
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, et al · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Earlier work this paper cites.
Learning fine-grained bimanual manipulation with low-cost hardware, 2023
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Earlier work this paper cites.
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Earlier work this paper cites.
Aligning large multimodal models with factually augmented rlhf, 2023
Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L.-Y. Gui, Y.-X. Wang, Y. Yang, K. Keutzer, and T. Darrell · 2023
Earlier work this paper cites.
Rh20t: A robotic dataset for learning diverse skills in one-shot
H.-S. Fang, H. Fang, Z. Tang, J. Liu, J. Wang, H. Zhu, and C. Lu · 2023
Earlier work this paper cites.
A system-level view on out-of-distribution data in robotics, 2023
R. Sinha, A. Sharma, S. Banerjee, T. Lew, R. Luo, S. M. Richards, Y. Sun, E. Schmerling, and M. Pavone · 2023
Earlier work this paper cites.
Code as policies: Language model programs for embodied control, 2023
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Cited alongside, same era.
Embodiedgpt: Vision-language pre-training via embodied chain of thought, 2023
Y. Mu, Q. Zhang, M. Hu, W. Wang, M. Ding, J. Jin, B. Wang, J. Dai, Y. Qiao, and P. Luo · 2023
Cited alongside, same era.
Libero: Benchmarking knowledge transfer for lifelong robot learning
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone · 2023
Cited alongside, same era.
π 0 \pi_{0} : A vision-language-action flow model for general robot control, 2024
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky · 2024
Cited alongside, same era.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al · 2024
Later among the works it cites.
Evaluating real-world robot manipulation policies in simulation
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, S. Levine, J. Wu, C. Finn, H. Su, Q. Vuong, and T. Xiao · 2024
Later among the works it cites.
Sglang: Efficient execution of structured language model programs, 2024
L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y. Sheng · 2024
Later among the works it cites.
Robotic control via embodied chain-of-thought reasoning, 2025
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Cited alongside, same era.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Cited alongside, same era.
Unpacking failure modes of generative policies: Runtime monitoring of consistency and progress, 2024
C. Agia, R. Sinha, J. Yang, Z. ang Cao, R. Antonova, M. Pavone, and J. Bohg · 2024
Cited alongside, same era.
Real-time anomaly detection and reactive planning with large language models, 2024
R. Sinha, A. Elhafsi, C. Agia, M. Foutter, E. Schmerling, and M. Pavone · 2024
Cited alongside, same era.
Autonomous improvement of instruction following skills via foundation models, 2024
Z. Zhou, P. Atreya, A. Lee, H. Walke, O. Mees, and S. Levine · 2024
Cited alongside, same era.
Re-mix: Optimizing data mixtures for large scale imitation learning, 2024
J. Hejna, C. Bhateja, Y. Jiang, K. Pertsch, and D. Sadigh · 2024
Cited alongside, same era.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
C.-L. Cheang, G. Chen, Y. Jing, T. Kong, H. Li, Y. Li, Y. Liu, H. Wu, J. Xu, Y. Yang, H. Zhang, and M. Zhu · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini · 2024
Cited alongside, same era.
J. Clark, S. Mirchandani, D. Sadigh, and S. Belkhale · 2025
Closest in time.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models, 2025
Q. Zhao, Y. Lu, M. J. Kim, Z. Fu, Z. Zhang, Y. Wu, Z. Li, Q. Ma, S. Han, C. Finn, A. Handa, M.-Y. Liu, D. Xiang, G. Wetzstein, and T.-Y. Lin · 2025
Closest in time.
Finetuning generative trajectory model with reinforcement learning from human feedback, 2025
D. Li, J. Ren, Y. Wang, X. Wen, P. Li, L. Xu, K. Zhan, Z. Xia, P. Jia, X. Lang, N. Xu, and H. Zhao · 2025
Closest in time.
S*: Test time scaling for code generation, 2025
D. Li, S. Cao, C. Cao, X. Li, S. Tan, K. Keutzer, J. Xing, J. E. Gonzalez, and I. Stoica · 2025
Closest in time.
How do large language monkeys get their power (laws)?, 2025
R. Schaeffer, J. Kazdan, J. Hughes, J. Juravsky, S. Price, A. Lynch, E. Jones, R. Kirk, A. Mirhoseini, and S. Koyejo · 2025
Closest in time.
Neural scaling laws in robotics, 2025
S. Sartor and N. Thompson · 2025
Closest in time.
Data scaling laws in imitation learning for robotic manipulation, 2025
F. Lin, Y. Hu, P. Sheng, C. Wen, J. You, and Y. Gao · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model, 2025
D. Qu, H. Song, Q. Chen, Y. Yao, X. Ye, Y. Ding, Z. Wang, J. Gu, B. Zhao, D. Wang, and X. Li · 2025
Closest in time.
Steering your generalists: Improving robotic foundation models via value guidance, 2025
M. Nakamoto, O. Mees, A. Kumar, and S. Levine · 2025
Closest in time.
Nerf-aug: Data augmentation for robotics with neural radiance fields, 2025
E. Zhu, M. Levy, M. Gwilliam, and A. Shrivastava · 2025
Closest in time.
Hi robot: Open-ended instruction following with hierarchical vision-language-action models, 2025
L. X. Shi, B. Ichter, M. Equi, L. Ke, K. Pertsch, Q. Vuong, J. Tanner, A. Walling, H. Wang, N. Fusai, A. Li-Bell, D. Driess, L. Groom, S. Levine, and C. Finn · 2025
Closest in time.
From foresight to forethought: Vlm-in-the-loop policy steering via latent alignment, 2025
Y. Wu, R. Tian, G. Swamy, and A. Bajcsy · 2025
Closest in time.
Inference-time policy steering through human interactions, 2025
Y. Wang, L. Wang, Y. Du, B. Sundaralingam, X. Yang, Y.-W. Chao, C. Perez-D’Arpino, D. Fox, and J. Shah · 2025
Closest in time.
Codemonkeys: Scaling test-time compute for software engineering, 2025
R. Ehrlich, B. Brown, J. Juravsky, R. Clark, C. Ré, and A. Mirhoseini · 2025
Closest in time.