Fetching the paper…
Reading the bibliography…
Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training.
E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. C. Courville, “Film: Visual reasoning with a general conditioning layer,” in AAAI . AAAI Press, 2018, pp. 3942–3951. [Online]. Available: https://doi.org/10.1609/aaai.v32i1.11671
2018
Earlier work this paper cites.
M. Weiler and G. Cesa, “General e(2)-equivariant steerable cnns,” in NIPS , 2019, pp. 14 334–14 345
2019
Earlier work this paper cites.
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine, “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in CoRL , ser. Proceedings of Machine Learning Research, vol. 100. PMLR, 2019, pp. 1094–1100
2019
Earlier work this paper cites.
Y. Zhou, C. Barnes, J. Lu, J. Yang, and H. Li, “On the continuity of rotation representations in neural networks,” in CVPR , 2019, pp. 5738–5746
2019
Earlier work this paper cites.
F. Fuchs, D. E. Worrall, V. Fischer, and M. Welling, “Se(3)-transformers: 3d roto-translation equivariant attention networks,” in NIPS , 2020
2020
Earlier work this paper cites.
V. G. Satorras, E. Hoogeboom, and M. Welling, “E(n) equivariant graph neural networks,” in ICML , ser. Proceedings of Machine Learning Research, vol. 139. PMLR, 2021, pp. 9323–9332
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu et al. , “RT-1: robotics transformer for real-world control at scale,” in RSS , 2023. [Online]. Available: https://doi.org/10.15607/RSS.2023.XIX.025
2023
Earlier work this paper cites.
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid et al. , “RT-2: vision-language-action models transfer web knowledge to robotic control,” in CoRL , ser. Proceedings of Machine Learning Research, vol. 229. PMLR, 2023, pp. 2165–2183
2023
Earlier work this paper cites.
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn, “Learning fine-grained bimanual manipulation with low-cost hardware,” in RSS , 2023. [Online]. Available: https://doi.org/10.15607/RSS.2023.XIX.016
2023
Earlier work this paper cites.
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” in RSS , 2023. [Online]. Available: https://doi.org/10.15607/RSS.2023.XIX.026
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in ICCV . IEEE, 2023, pp. 11 941–11 952. [Online]. Available: https://doi.org/10.1109/ICCV51070.2023.01100
2023
Earlier work this paper cites.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. C. H. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” in NeurIPS , 2023
2023
Cited alongside, same era.
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone, “LIBERO: benchmarking knowledge transfer for lifelong robot learning,” in NeurIPS , 2023
2023
Cited alongside, same era.
2024
Cited alongside, same era.
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong et al. , “Openvla: An open-source vision-language-action model,” in CoRL , ser. Proceedings of Machine Learning Research, vol. 270. PMLR, 2024, pp. 2679–2713
2024
Cited alongside, same era.
M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al. , “Dinov2: Learning robust visual features without supervision,” Trans. Mach. Learn. Res. , vol. 2024, 2024
2024
Later among the works it cites.
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis et al. , “DROID: A large-scale in-the-wild robot manipulation dataset,” in RSS , 2024. [Online]. Available: https://doi.org/10.15607/RSS.2024.XX.120
2024
Later among the works it cites.
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu, “RDT-1B: a diffusion foundation model for bimanual manipulation,” in ICLR . OpenReview.net, 2025
2025
Closest in time.
2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Ghosh, H. R. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo et al. , “Octo: An open-source generalist robot policy,” in RSS , 2024. [Online]. Available: https://doi.org/10.15607/RSS.2024.XX.090
2024
Cited alongside, same era.
H. Wu, Y. Jing, C. Cheang, G. Chen, J. Xu, X. Li, M. Liu, H. Li, and T. Kong, “Unleashing large-scale video generative pre-training for visual robot manipulation,” in ICLR . OpenReview.net, 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
C. Wen, X. Lin, J. I. R. So, K. Chen, Q. Dou, Y. Gao, and P. Abbeel, “Any-point trajectory modeling for policy learning,” in RSS , 2024. [Online]. Available: https://doi.org/10.15607/RSS.2024.XX.092
2024
Cited alongside, same era.
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain et al. , “Open x-embodiment: Robotic learning datasets and RT-X models : Open x-embodiment collaboration,” in ICRA . IEEE, 2024, pp. 6892–6903. [Online]. Available: https://doi.org/10.1109/ICRA57147.2024.10611477
2024
Cited alongside, same era.
J. H. Yang, C. Glossop, A. Bhorkar, D. Shah, Q. Vuong, C. Finn, D. Sadigh, and S. Levine, “Pushing the limits of cross-embodiment learning for manipulation and navigation,” in RSS , 2024. [Online]. Available: https://doi.org/10.15607/RSS.2024.XX.093
2024
Cited alongside, same era.
C. Wang, H. Fang, H.-S. Fang, and C. Lu, “Rise: 3d perception makes real-world robot imitation simple and effective,” in IROS , 2024, pp. 2870–2877
2024
Cited alongside, same era.
T. Miyato, B. Jaeger, M. Welling, and A. Geiger, “GTA: A geometry-aware attention mechanism for multi-view transformers,” in ICLR , 2024
2024
Cited alongside, same era.
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” in CVPR , 2025
2025
Closest in time.