Fetching the paper…
Reading the bibliography…
Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
Jang, E., Irpan, A., Khansari, M., Kappler, D., Ebert, F., Lynch, C., Levine, S., and Finn, C · 2022
Earlier work this paper cites.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Mees, O., Hermann, L., Rosete-Beas, E., and Burgard, W · 2022
Earlier work this paper cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., et al · 2022
Earlier work this paper cites.
Zero-shot robotic manipulation with pretrained image-editing diffusion models
Black, K., Nakamoto, M., Atreya, P., Walke, H., Finn, C., Kumar, A., and Levine, S · 2023
Earlier work this paper cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models
O’Neill, A., Rehman, A., Gupta, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., Pooley, A., Gupta, A., Mandlekar, A., et al · 2023
Cited alongside, same era.
Bridgedata v2: A dataset for robot learning at scale
Seed-x: Multimodal models with unified multi-granularity comprehension and generation
Ge, Y., Zhao, S., Zhu, J., Ge, Y., Yi, K., Song, L., Li, C., Ding, X., and Shan, Y · 2024
Later among the works it cites.
Exploring failure cases in multimodal reasoning about physical dynamics
Ghaffari, S. and Krishnaswamy, N · 2024
Later among the works it cites.
Prediction with action: Visual policy learning via joint denoising process
Guo, Y., Hu, Y., Zhang, J., Wang, Y.-J., Chen, X., Lu, C., and Chen, J · 2024
Later among the works it cites.
3d diffuser actor: Policy diffusion with 3d scene representations
Ke, T.-W., Gkanatsios, N., and Fragkiadaki, K · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Walke, H., Black, K., Lee, A., Kim, M. J., Du, M., Zheng, C., Zhao, T., Hansen-Estruch, P., Vuong, Q., He, A., Myers, V., Fang, K., Finn, C., and Levine, S · 2023
Cited alongside, same era.
Prompt a robot to walk with large language models
Wang, Y.-J., Zhang, B., Chen, J., and Sreenath, K · 2023
Cited alongside, same era.
Unleashing large-scale video generative pre-training for visual robot manipulation
Wu, H., Jing, Y., Cheang, C., Chen, G., Xu, J., Li, X., Liu, M., Li, H., and Kong, T · 2023
Cited alongside, same era.
Synthetic vision: Training vision-language models to understand physics
Balazadeh, V., Ataei, M., Cheong, H., Khasahmadi, A. H., and Krishnan, R. G · 2024
Cited alongside, same era.
p i _ 0 pi\_0 : A vision-language-action flow model for general robot control
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al · 2024
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P. N., and Hoi, S · 2024
Cited alongside, same era.
Learning universal policies via text-guided video generation
Du, Y., Yang, S., Dai, B., Dai, H., Nachum, O., Tenenbaum, J., Schuurmans, D., and Abbeel, P · 2024
Cited alongside, same era.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Chen, B., Xu, Z., Kirmani, S., Ichter, B., Sadigh, D., Guibas, L., and Xia, F
Cited in the paper.
Kim, M., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., Vuong, Q., Kollar, T., Burchfiel, B., Tedrake, R., Sadigh, D., Levine, S., Liang, P., and Finn, C · 2024
Later among the works it cites.
Li, Z., Wang, H., Liu, D., Zhang, C., Ma, A., Long, J., and Cai, W · 2024
Later among the works it cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2024
Later among the works it cites.
Can transformers capture spatial relations between objects?
Wen, C., Jayaraman, D., and Gao, Y · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z · 2024
Later among the works it cites.
Hirt: Enhancing robotic control with hierarchical robot transformers
Zhang, J., Guo, Y., Chen, X., Wang, Y.-J., Hu, Y., Shi, C., and Chen, J · 2024
Later among the works it cites.
3d-vla: A 3d vision-language-action generative world model
Zhen, H., Qiu, X., Chen, P., Yang, J., Yan, X., Du, Y., Hong, Y., and Gan, C · 2024
Later among the works it cites.