Fetching the paper…
Reading the bibliography…
Vision-language-action (VLA) models have significantly advanced robotic learning, enabling training on large-scale, cross-embodiment data and fine-tuning for specific robots.
Rlbench: The robot learning benchmark & learning environment
James, S., Ma, Z., Arrojo, D. R., and Davison, A. J · 2020
Earlier work this paper cites.
Unilog: Deploy one model and specialize it for all log analysis tasks
Zhu, Y., Meng, W., Liu, Y., Zhang, S., Han, T., Tao, S., and Pei, D · 2021
Earlier work this paper cites.
Label-guided auxiliary training improves 3d object detector
Huang, Y., Liu, X., Zhu, Y., Xu, Z., Shen, C., Che, Z., Zhang, G., Peng, Y., Feng, F., and Tang, J · 2022
Earlier work this paper cites.
Vima: General robot manipulation with multimodal prompts
Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L · 2022
Earlier work this paper cites.
Teach less, learn more: On the undistillable classes in knowledge distillation
Zhu, Y., Liu, N., Xu, Z., Liu, X., Meng, W., Wang, L., Ou, Z., and Tang, J · 2022
Earlier work this paper cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Awadalla, A., Gao, I., Gardner, J., Hessel, J., Hanafy, Y., Zhu, W., Marathe, K., Bitton, Y., Gadre, S., Sagawa, S., et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Earlier work this paper cites.
Rvt: Robotic view transformer for 3d object manipulation
Goyal, A., Xu, J., Guo, Y., Blukis, V., Chao, Y.-W., and Fox, D · 2023
Earlier work this paper cites.
Obelisc: An open web-scale filtered dataset of interleaved image-text documents
Laurençon, H., Saulnier, L., Tronchon, L., Bekman, S., Singh, A., Lozhkov, A., Wang, T., Karamcheti, S., Rush, A., and Kiela, D · 2023
Earlier work this paper cites.
Vision-language foundation models as effective robot imitators
Li, X., Liu, M., Zhang, H., Yu, C., Xu, J., Wu, H., Cheang, C., Jing, Y., Zhang, W., Liu, H., et al · 2023
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Earlier work this paper cites.
Logsummary: Unstructured log summarization for software systems
Meng, W., Zaiter, F., Zhang, Y., Liu, Y., Zhang, S., Tao, S., Zhu, Y., Han, T., Zhao, Y., Wang, E., et al · 2023
Earlier work this paper cites.
Sayplan: Grounding large language models using 3d scene graphs for scalable task planning
Rana, K., Haviland, J., Garg, S., Abou-Chakra, J., Reid, I. D., and Suenderhauf, N · 2023
Earlier work this paper cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2023
Earlier work this paper cites.
Bridgedata v2: A dataset for robot learning at scale
Walke, H. R., Black, K., Zhao, T. Z., Vuong, Q., Zheng, C., Hansen-Estruch, P., He, A. W., Myers, V., Kim, M. J., Du, M., et al · 2023
Cited alongside, same era.
Make a long image short: Adaptive token length for vision transformers
Zhou, Q. and Zhu, Y · 2023
Cited alongside, same era.
π 0 \pi_{0} : A vision-language-action flow model for general robot control, 2024
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., Jakubczak, S., Jones, T., Ke, L., Levine, S., Li-Bell, A., Mothukuri, M., Nair, S., Pertsch, K., Shi, L. X., Tanner, J., Vuong, Q., Walling, A., Wang, H., and Zhilinsky, U · 2024
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., et al · 2024
Cited alongside, same era.
Robotwin: Dual-arm robot benchmark with generative digital twins (early version)
Mu, Y., Chen, T., Peng, S., Chen, Z., Gao, Z., Zou, Y., Lin, L., Xie, Z., and Luo, P · 2024
Later among the works it cites.
Llarva: Vision-action instruction tuning enhances robot learning
Niu, D., Sharma, Y., Biamby, G., Quenum, J., Bai, Y., Shi, B., Darrell, T., and Herzig, R · 2024
Later among the works it cites.
Consistency policy: Accelerated visuomotor policies via consistency distillation
Prasad, A., Lin, K., Wu, J., Zhou, L., and Bohg, J · 2024
Later among the works it cites.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
Reuss, M., Yağmurlu, Ö. E., Wenzel, F., and Lioutikov, R · 2024
Later among the works it cites.
Object-centric instruction augmentation for robotic manipulation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dasari, S., Mees, O., Zhao, S., Srirama, M. K., and Levine, S · 2024
Cited alongside, same era.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Fu, Z., Zhao, T. Z., and Finn, C · 2024
Cited alongside, same era.
Dag-plan: Generating directed acyclic dependency graphs for dual-arm cooperative planning
Gao, Z., Mu, Y., Qu, J., Hu, M., Guo, L., Luo, P., and Lu, Y · 2024
Cited alongside, same era.
Peract2: Benchmarking and learning for robotic bimanual manipulation tasks
Grotz, M., Shridhar, M., Chao, Y.-W., Asfour, T., and Fox, D · 2024
Cited alongside, same era.
An embodied generalist agent in 3d world
Huang, J., Yong, S., Ma, X., Linghu, X., Li, P., Wang, Y., Li, Q., Zhu, S.-C., Jia, B., and Huang, S · 2024
Cited alongside, same era.
Ragraph: A general retrieval-augmented graph learning framework
Jiang, X., Qiu, R., Xu, Y., Zhu, Y., Zhang, R., Fang, Y., Xu, C., Zhao, J., and Wang, Y · 2024
Cited alongside, same era.
Prismatic vlms: Investigating the design space of visually-conditioned language models
Karamcheti, S., Nair, S., Balakrishna, A., Liang, P., Kollar, T., and Sadigh, D · 2024
Cited alongside, same era.
Droid: A large-scale in-the-wild robot manipulation dataset
Khazatsky, A., Pertsch, K., Nair, S., Balakrishna, A., Dasari, S., Karamcheti, S., Nasiriany, S., Srirama, M. K., Chen, L. Y., Ellis, K., et al · 2024
Cited alongside, same era.
Wen, J., Zhu, Y., Zhu, M., Li, J., Xu, Z., Che, Z., Shen, C., Peng, Y., Liu, D., Feng, F., and Tang, J · 2024
Later among the works it cites.
Object-centric instruction augmentation for robotic manipulation
Wen, J., Zhu, Y., Zhu, M., Li, J., Xu, Z., Che, Z., Shen, C., Peng, Y., Liu, D., Feng, F., et al · 2024
Later among the works it cites.
Robotic control via embodied chain-of-thought reasoning
Zawalski, M., Chen, W., Pertsch, K., Mees, O., Finn, C., and Levine, S · 2024
Later among the works it cites.
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models, 2024
Zhao, Q., Lu, Y., Kim, M. J., Fu, Z., Zhang, Z., Wu, Y., Ma, M. L. Q., Han, S., Finn, C., Handa, A., Liu, M.-Y., Xiang, D., Wetzstein, G., and Lin, T.-Y · 2024
Later among the works it cites.
Language-conditioned robotic manipulation with fast and slow thinking
Zhu, M., Zhu, Y., Li, J., Wen, J., Xu, Z., Che, Z., Shen, C., Peng, Y., Liu, D., Feng, F., and Tang, J · 2024
Later among the works it cites.
Language-conditioned robotic manipulation with fast and slow thinking
Zhu, M., Zhu, Y., Li, J., Wen, J., Xu, Z., Che, Z., Shen, C., Peng, Y., Liu, D., Feng, F., et al · 2024
Later among the works it cites.
Mm-safetybench: A benchmark for safety evaluation of multimodal large language models
Liu, X., Zhu, Y., Gu, J., Lan, Y., Yang, C., and Qiao, Y · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
Pertsch, K., Stachowicz, K., Ichter, B., Driess, D., Nair, S., Vuong, Q., Mees, O., Finn, C., and Levine, S · 2025
Closest in time.
Discrete policy: Learning disentangled action space for multi-task robotic manipulation
Wu, K., Zhu, Y., Li, J., Wen, J., Liu, N., Xu, Z., and Tang, J · 2025
Closest in time.
Vision-language-action model with open-world embodied reasoning from pretrained knowledge
Zhou, Z., Zhu, Y., Wen, J., Shen, C., and Xu, Y · 2025
Closest in time.
Scaling diffusion policy in transformer to 1 billion parameters for robotic manipulation
Zhu, M., Zhu, Y., Li, J., Wen, J., Xu, Z., Liu, N., Cheng, R., Shen, C., Peng, Y., Feng, F., et al · 2025
Closest in time.