Fetching the paper…
Reading the bibliography…
Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes.
E. Perez, F. Strub et al. , “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
T. Yu, D. Quillen et al. , “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in Conference on robot learning . PMLR, 2020, pp. 1094–1100
2020
Earlier work this paper cites.
S. Toyer, R. Shah, A. Critch, and S. Russell, “The magical benchmark for robust imitation,” Advances in Neural Information Processing Systems , vol. 33, pp. 18 284–18 295, 2020
2020
Earlier work this paper cites.
D. Yarats, I. Kostrikov, and R. Fergus, “Image augmentation is all you need: Regularizing deep reinforcement learning from pixels,” in International conference on learning representations , 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
F. Ebert, Y. Yang et al. , “Bridge data: Boosting generalization of robotic skills with cross-domain datasets,” RSS , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. Jang, A. Irpan et al. , “Bc-z: Zero-shot task generalization with robotic imitation learning,” in Conference on Robot Learning . PMLR, 2022, pp. 991–1002
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. J. Hu, yelong shen et al. , “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
J. Lu, C. Clark, R. Zellers, R. Mottaghi, and A. Kembhavi, “Unified-io: A unified model for vision, language, and multi-modal tasks,” in The Eleventh International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
T. Chen, S. Saxena, L. Li, T.-Y. Lin, D. J. Fleet, and G. E. Hinton, “A unified sequence interface for vision tasks,” Advances in Neural Information Processing Systems , vol. 35, pp. 31 333–31 346, 2022
2022
Earlier work this paper cites.
L. A. Doumas, G. Puebla, A. E. Martin, and J. E. Hummel, “A theory of relation learning and cross-domain generalization.” Psychological review , vol. 129, no. 5, p. 999, 2022
2022
Earlier work this paper cites.
C. Chi, S. Feng et al. , “Diffusion policy: Visuomotor policy learning via action diffusion,” RSS , 2023
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Ze, G. Zhang et al. , “3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations,” in Proceedings of Robotics: Science and Systems (RSS) , 2024
2024
Closest in time.
M. J. Kim, K. Pertsch et al. , “Openvla: An open-source vision-language-action model,” 8th Annual Conference on Robot Learning , 2024
2024
Closest in time.
D. Zhu, J. Chen et al. , “MiniGPT-4: Enhancing vision-language understanding with advanced large language models,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
K. Rana, J. Haviland et al. , “Sayplan: Grounding large language models using 3d scene graphs for scalable task planning,” in 7th Annual Conference on Robot Learning , 2023
2023
Cited alongside, same era.
S. Biderman, H. Schoelkopf et al. , “Pythia: A suite for analyzing large language models across training and scaling,” in International Conference on Machine Learning . PMLR, 2023, pp. 2397–2430
2023
Cited alongside, same era.
T. Chen, L. Li et al. , “A generalist framework for panoptic segmentation of images and videos,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 909–919
2023
Cited alongside, same era.
2024
Closest in time.
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine, “Robotic control via embodied chain-of-thought reasoning,” in Conference on Robot Learning (CoRL) , vol. 270, 2024, pp. 3157–3181
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Ze, Y. Liu et al. , “H-index: Visual reinforcement learning with hand-informed representations for dexterous manipulation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
Octo Model Team, D. Ghosh, H. Walke et al. , “Octo: An open-source generalist robot policy,” in Proceedings of Robotics: Science and Systems , Delft, Netherlands, 2024
2024
Closest in time.
M. Reuss, Ö. E. Yağmurlu, F. Wenzel, and R. Lioutikov, “Multimodal diffusion transformer: Learning versatile behavior from multimodal goals,” Robotics: Science and Systems , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Z. Zhao, J. Tompson et al. , “Aloha unleashed: A simple recipe for robot dexterity,” in 8th Annual Conference on Robot Learning , 2024
2024
Closest in time.