Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models have achieved notable success but often struggle with limited generalizations.
Virtualhome: Simulating household activities via programs
Puig, X., Ra, K., Boben, M., Li, J., Wang, T., Fidler, S., and Torralba, A · 2018
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations
Guss, W. H., Houghton, B., Topin, N., Wang, P., Codel, C., Veloso, M., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettlemoyer, L., and Fox, D · 2020
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Earlier work this paper cites.
Camel: Communicative agents for” mind” exploration of large language model society
Li, G., Hammoud, H., Itani, H., Khizbullin, D., and Ghanem, B · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Earlier work this paper cites.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Song, C. H., Wu, J., Washington, C., Sadler, B. M., Chao, W.-L., and Su, Y · 2023
Earlier work this paper cites.
Paligemma: A versatile 3b vlm for transfer
Beyer, L., Steiner, A., Pinto, A. S., Kolesnikov, A., Wang, X., Salz, D., Neumann, M., Alabdulmohsin, I., Tschannen, M., Bugliarello, E., et al · 2024
Earlier work this paper cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al · 2024
Earlier work this paper cites.
Scaling robot policy learning via zero-shot labeling with foundation models
Blank, N., Reuss, M., Rühle, M., Yağmurlu, Ö. E., Wenzel, F., Mees, O., and Lioutikov, R · 2024
Earlier work this paper cites.
Genie: Generative interactive environments
Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al · 2024
Earlier work this paper cites.
Towards synergistic, generalized, and efficient dual-system for robotic manipulation
Bu, Q., Li, H., Chen, L., Cai, J., Zeng, J., Cui, H., Yao, M., and Qiao, Y · 2024
Cited alongside, same era.
Rovi-aug: Robot and viewpoint augmentation for cross-embodiment robot learning
Chen, L. Y., Xu, C., Dharmarajan, K., Irshad, M. Z., Cheng, R., Keutzer, K., Tomizuka, M., Vuong, Q., and Goldberg, K · 2024
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S · 2024
Cited alongside, same era.
Gemini 2.0 flash thinking
DeepMind, G · 2024
Cited alongside, same era.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Cited alongside, same era.
Oasis: Open agents social interaction simulations on one million agents
Yang, Z., Zhang, Z., Zheng, Z., Jiang, Y., Gan, Z., Wang, Z., Ling, Z., Chen, J., Ma, M., Dong, B., et al · 2024
Later among the works it cites.
Gr00t n1: An open foundation model for generalist humanoid robots
Bjorck, J., Castañeda, F., Cherniadev, N., Da, X., Ding, R., Fan, L., Fang, Y., Fox, D., Hu, F., Huang, S., et al · 2025
Closest in time.
Conrft: A reinforced fine-tuning method for vla models via consistency policy
Chen, Y., Tian, S., Zhou, Y., Liu, S., Li, H., and Zhao, D · 2025
Closest in time.
π 0.5 \pi_{0.5} : a vision-language-action model with open-world generalization, 2025
Intelligence, P., Black, K., Brown, N., Darpinian, J., Dhabalia, K., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Galliker, M. Y., Ghosh, D., Groom, L., Hausman, K., Ichter, B., Jakubczak, S., Jones, T., Ke, L., LeBlanc, D., Levine, S., Li-Bell, A., Mothukuri, M., Nair, S., Pertsch, K., Ren, A. Z., Shi, L. X., Smith, L., Springenberg, J. T., Stachowicz, K., Tanner, J., Vuong, Q., Walke, H., Walling, A., Wang, H., Yu, L., and Zhilinsky, U · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Openvla: An open-source vision-language-action model
Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al · 2024
Cited alongside, same era.
Behavior generation with latent actions
Lee, S., Wang, Y., Etukuru, H., Kim, H. J., Shafiullah, N. M. M., and Pinto, L · 2024
Cited alongside, same era.
Data scaling laws in imitation learning for robotic manipulation
Lin, F., Hu, Y., Sheng, P., Wen, C., You, J., and Gao, Y · 2024
Cited alongside, same era.
A survey on vision-language-action models for embodied ai
Ma, Y., Song, Z., Zhuang, Y., Hao, J., and King, I · 2024
Cited alongside, same era.
Nguyen, D., Chen, J., Wang, Y., Wu, G., Park, N., Hu, Z., Lyu, H., Wu, J., Aponte, R., Xia, Y., et al · 2024
Cited alongside, same era.
Octo: An open-source generalist robot policy
Octo Model Team, Ghosh, D., Walke, H., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Xu, C., Luo, J., Kreiman, T., Tan, Y., Chen, L. Y., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., and Levine, S · 2024
Cited alongside, same era.
Mp5: A multi-modal open-ended embodied system in minecraft via active perception
Qin, Y., Zhou, E., Liu, Q., Yin, Z., Sheng, L., Zhang, R., Qiao, Y., and Shao, J · 2024
Cited alongside, same era.
Closest in time.
Vision language models are in-context value learners
Ma, Y. J., Hejna, J., Fu, C., Shah, D., Liang, J., Xu, Z., Kirmani, S., Xu, P., Driess, D., Xiao, T., Bastani, O., Jayaraman, D., Yu, W., Zhang, T., Sadigh, D., and Xia, F · 2025
Closest in time.
Sim-and-real co-training: A simple recipe for vision-based robotic manipulation
Maddukuri, A., Jiang, Z., Chen, L. Y., Nasiriany, S., Xie, Y., Fang, Y., Huang, W., Wang, Z., Xu, Z., Chernyadev, N., et al · 2025
Closest in time.
Flower: Democratizing generalist robot policies with efficient vision-language-action flow policies
Reuss, M., Zhou, H., Rühle, M., Yağmurlu, Ö. E., Otto, F., and Lioutikov, R · 2025
Closest in time.
Hi robot: Open-ended instruction following with hierarchical vision-language-action models
Shi, L. X., Ichter, B., Equi, M., Ke, L., Pertsch, K., Vuong, Q., Tanner, J., Walling, A., Wang, H., Fusai, N., et al · 2025
Closest in time.
Doubao-vision-pro-32k
Team, D · 2025
Closest in time.
The rise and potential of large language model based agents: A survey
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., et al · 2025
Closest in time.
Latent action pretraining from videos
Ye, S., Jang, J., Jeon, B., Joo, S. J., Yang, J., Peng, B., Mandlekar, A., Tan, R., Chao, Y.-W., Lin, B. Y., Liden, L., Lee, K., Gao, J., Zettlemoyer, L., Fox, D., and Seo, M · 2025
Closest in time.
Universal actions for enhanced embodied foundation models
Zheng, J., Li, J., Liu, D., Zheng, Y., Wang, Z., Ou, Z., Liu, Y., Liu, J., Zhang, Y.-Q., and Zhan, X · 2025
Closest in time.
Dexgraspvla: A vision-language-action framework towards general dexterous grasping
Zhong, Y., Huang, X., Li, R., Zhang, C., Liang, Y., Yang, Y., and Chen, Y · 2025
Closest in time.