With their prominent scene understanding and reasoning capabilities, pre-trained visual-language models (VLMs) such as GPT-4V have attracted increasing attention in robotic task planning.
Compared with traditional task planning strategies, VLMs are strong in multimodal information parsing and code generation and show remarkable efficiency.
Although VLMs demonstrate great potential in robotic task planning, they suffer from challenges like hallucination, semantic complexity, and limited context.
To handle such issues, this paper proposes a multi-agent framework, i.e., GameVLM, to enhance the decision-making process in robotic task planning.
GameVLM: A Decision-making Framework for Robotic Task Planning Based on Visual Language Models and Zero-sum Games · Around
Built on
R. Alterovitz, S. Koenig, and M. Likhachev, “Robot planning in the real world: Research challenges and opportunities,” AI Magazine , vol. 37, no. 2, pp. 76–84, 2016
2016
Earlier work this paper cites.
S.-H. Hwang and L. Rey-Bellet, “Strategic decompositions of normal form games: Zero-sum games and potential games,” Games and Economic Behavior , vol. 122, pp. 370–390, 2020
Z. Bao, G.-N. Zhu, W. Ding, Y. Guan, W. Bai, and Z. Gan, “A smart interactive camera robot based on large language models,” in 2023 IEEE International Conference on Robotics and Biomimetics (ROBIO) . IEEE, 2023, pp. 1–6
2023
Cited alongside, same era.
D. Shah, B. Osiński, S. Levine et al. , “LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action,” in 6th Conference on Robot Learning (CoRL 2022) . PMLR, 2023, pp. 492–504
R. Melnic, V. Ababii, V. Sudacevschi, O. Sachenko, O. Borozan, and T. Lendiuk, “Multi-objective based multi-agent decision-making system,” in 2023 IEEE 12th International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications (IDAACS) , vol. 1. IEEE, 2023, pp. 834–839