Fetching the paper…
Reading the bibliography…
Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks.
V-rep: A versatile and scalable robot simulation framework
Rohmer, E., Singh, S. P., and Freese, M · 2013
Earlier work this paper cites.
The ycb object and model set: Towards common benchmarks for manipulation research
Calli, B., Singh, A., Walsman, A., Srinivasa, S., Abbeel, P., and Dollar, A. M · 2015
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Redmon, J · 2016
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Deitke, M., Ehsani, K., Gordon, D., Zhu, Y., et al · 2017
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs
Puig, X., Ra, K., Boben, M., Li, J., Wang, T., Fidler, S., and Torralba, A · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
James, S., Ma, Z., Arrojo, D. R., and Davison, A. J · 2020
Earlier work this paper cites.
Sapien: A simulated part-based interactive environment
Xiang, F., Qin, Y., Mo, K., Xia, Y., Zhu, H., Liu, F., Liu, M., Jiang, H., Yuan, Y., Wang, H., et al · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Robustnav: Towards benchmarking robustness in embodied navigation
Chattopadhyay, P., Hoffman, J., Mottaghi, R., and Kembhavi, A · 2021
Earlier work this paper cites.
igibson 2.0: Object-centric simulation for robot learning of everyday household tasks
Li, C., Xia, F., Martín-Martín, R., Lingelbach, M., Srivastava, S., Shen, B., Vainio, K., Gokmen, C., Dharan, G., Jain, T., et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
igibson 1.0: A simulation environment for interactive tasks in large realistic scenes
Shen, B., Xia, F., Li, C., Martín-Martín, R., Fan, L., Wang, G., Pérez-D’Arpino, C., Buch, S., Srivastava, S., Tchapmi, L., et al · 2021
Earlier work this paper cites.
Habitat 2.0: Training home assistants to rearrange their habitat
Szot, A., Clegg, A., Undersander, E., Wijmans, E., Zhao, Y., Turner, J., Maestre, N., Mukadam, M., Chaplot, D. S., Maksymets, O., et al · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2022
Earlier work this paper cites.
Robotic skill acquisition via instruction augmentation with vision-language models
Xiao, T., Chan, H., Sermanet, P., Wahid, A., Brohan, A., Hausman, K., Levine, S., and Tompson, J · 2022
Earlier work this paper cites.
Vlmbench: A compositional benchmark for vision-and-language manipulation
Zheng, K., Chen, X., Jenkins, O. C., and Wang, X · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Compositional foundation models for hierarchical planning
Ajay, A., Han, S., Du, Y., Li, S., Gupta, A., Jaakkola, T., Tenenbaum, J., Kaelbling, L., Srivastava, A., and Agrawal, P · 2023
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Earlier work this paper cites.
Vistruct: Visual structural knowledge extraction via curriculum guided code-vision representation
Chen, Y., Wang, X., Li, M., Hoiem, D., and Ji, H · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S · 2023
Earlier work this paper cites.
Lmdeploy: A toolkit for compressing, deploying, and serving llm
Contributors, L · 2023
Earlier work this paper cites.
Palm-e: an embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Earlier work this paper cites.
Du, Y., Yang, M., Florence, P., Xia, F., Wahid, A., Ichter, B., Sermanet, P., Yu, T., Abbeel, P., Tenenbaum, J. B., et al · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation
Li, C., Zhang, R., Wong, J., Gokmen, C., Srivastava, S., Martín-Martín, R., Wang, C., Levine, G., Lingelbach, M., Sun, J., et al · 2023
Cited alongside, same era.
Code as policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Cited alongside, same era.
Fmb: a functional manipulation benchmark for generalizable robotic learning
Luo, J., Xu, C., Liu, F., Tan, L., Lin, Z., Wu, J., Abbeel, P., and Levine, S · 2023
Cited alongside, same era.
Gpt-driver: Learning to drive with gpt
Mao, J., Qian, Y., Zhao, H., and Wang, Y · 2023
Cited alongside, same era.
Sayplan: Grounding large language models using 3d scene graphs for scalable task planning
Rana, K., Haviland, J., Garg, S., Abou-Chakra, J., Reid, I. D., and Suenderhauf, N · 2023
Cited alongside, same era.
Openvla: An open-source vision-language-action model
Kim, M. J., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., et al · 2024
Later among the works it cites.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., and Fried, D · 2024
Later among the works it cites.
Muep: A multimodal benchmark for embodied planning with foundation models
Li, K., Yu, B., Zheng, Q., Zhan, Y., Zhang, Y., Zhang, T., Yang, Y., Chen, Y., Sun, L., Cao, Q., Shen, L., Li, L., Tao, D., and He, X · 2024
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Later among the works it cites.
Ovis: Structural embedding alignment for multimodal large language model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semantic mechanical search with large vision and language models
Sharma, S., Huang, H., Shivakumar, K., Chen, L. Y., Hoque, R., Ichter, B., and Goldberg, K · 2023
Cited alongside, same era.
Progprompt: Generating situated robot task plans using large language models
Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A · 2023
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Song, C. H., Wu, J., Washington, C., Sadler, B. M., Chao, W.-L., and Su, Y · 2023
Cited alongside, same era.
Open-world object manipulation using pre-trained vision-language models
Stone, A., Xiao, T., Lu, Y., Gopalakrishnan, K., Lee, K.-H., Vuong, Q., Wohlhart, P., Kirmani, S., Zitkovich, B., Xia, F., et al · 2023
Cited alongside, same era.
Sun, H · 2023
Cited alongside, same era.
Large language models as generalizable policies for embodied tasks
Szot, A., Schwarzer, M., Agrawal, H., Mazoure, B., Metcalf, R., Talbott, W., Mackraz, N., Hjelm, R. D., and Toshev, A. T · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Lu, S., Li, Y., Chen, Q.-G., Xu, Z., Luo, W., Zhang, K., and Ye, H.-J · 2024
Later among the works it cites.
Foundation models for video understanding: A survey
Madan, N., Møgelmose, A., Modi, R., Rawat, Y. S., and Moeslund, T. B · 2024
Later among the works it cites.
Genrl: Multimodal-foundation world models for generalization in embodied agents
Mazzaglia, P., Verbelen, T., Dhoedt, B., Courville, A., and Rajeswar, S · 2024
Later among the works it cites.
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024
Meta · 2024
Later among the works it cites.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Mu, Y., Zhang, Q., Hu, M., Wang, W., Ding, M., Jin, J., Wang, B., Dai, J., Qiao, Y., and Luo, P · 2024
Later among the works it cites.
Escapebench: Pushing language models to think outside the box
Qian, C., Han, P., Luo, Q., He, B., Chen, X., Zhang, Y., Du, H., Yao, J., Yang, X., Zhang, D., Li, Y., and Ji, H · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Later among the works it cites.
Towards long-horizon vision-language navigation: Platform, benchmark and method
Song, X., Chen, W., Liu, Y., Chen, W., Li, G., and Lin, L · 2024
Later among the works it cites.
From multimodal llms to generalist embodied agents: Methods and lessons
Szot, A., Mazoure, B., Attia, O., Timofeev, A., Agrawal, H., Hjelm, D., Gan, Z., Kira, Z., and Toshev, A · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., et al · 2024
Later among the works it cites.
Dissecting adversarial robustness of multimodal lm agents
Wu, C. H., Shah, R. R., Koh, J. Y., Salakhutdinov, R., Fried, D., and Raghunathan, A · 2024
Later among the works it cites.
Pandora: Towards general world model with natural language actions and video states
Xiang, J., Liu, G., Gu, Y., Gao, Q., Ning, Y., Zha, Y., Feng, Z., Tao, T., Hao, S., Shi, Y., et al · 2024
Later among the works it cites.
Large multimodal agents: A survey
Xie, J., Chen, Z., Zhang, R., Wan, X., and Li, G · 2024
Later among the works it cites.
Robust decision transformer: Tackling data corruption in offline rl via sequence modeling
Xu, J., Yang, R., Luo, F., Fang, M., Wang, B., and Han, L · 2024
Later among the works it cites.
In-context learning enables robot action prediction in llms
Yin, Y., Wang, Z., Sharma, Y., Niu, D., Darrell, T., and Herzig, R · 2024
Later among the works it cites.
Robotic control via embodied chain-of-thought reasoning
Zawalski, M., Chen, W., Pertsch, K., Mees, O., Finn, C., and Levine, S · 2024
Later among the works it cites.
Cosmos world foundation model platform for physical ai
Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al · 2025
Closest in time.
Qwen2.5-vl technical report, 2025
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., and Lin, J · 2025
Closest in time.
Chen, Z., Wang, W., Cao, Y., Liu, Y., Gao, Z., Cui, E., Zhu, J., Ye, S., Tian, H., Liu, Z., Gu, L., Wang, X., Li, Q., Ren, Y., Chen, Z., Luo, J., Wang, J., Jiang, T., Wang, B., He, C., Shi, B., Zhang, X., Lv, H., Wang, Y., Shao, W., Chu, P., Tu, Z., He, T., Wu, Z., Deng, H., Ge, J., Chen, K., Zhang, K., Wang, L., Dou, M., Lu, L., Zhu, X., Lu, T., Lin, D., Qiao, Y., Dai, J., and Wang, W · 2025
Closest in time.
Embodiedeval: Evaluate multimodal llms as embodied agents
Cheng, Z., Tu, Y., Li, R., Dai, S., Hu, J., Hu, S., Li, J., Shi, Y., Yu, T., Chen, W., et al · 2025
Closest in time.
Kimi k1.5: Scaling reinforcement learning with llms
Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., Tang, C., Wang, C., Zhang, D., Yuan, E., Lu, E., Tang, F., Sung, F., Wei, G., Lai, G., Guo, H., Zhu, H., et al · 2025
Closest in time.
Physgen: Rigid-body physics-grounded image-to-video generation
Liu, S., Ren, Z., Gupta, S., and Wang, S · 2025
Closest in time.
Team, G., Kamath, A., Ferret, J., Pathak, S., Vieillard, N., Merhej, R., Perrin, S., Matejovicova, T., Ramé, A., Rivière, M., et al · 2025
Closest in time.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Zhai, S., Bai, H., Lin, Z., Pan, J., Tong, P., Zhou, Y., Suhr, A., Xie, S., LeCun, Y., Ma, Y., et al · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models
Zhu, J., Wang, W., Chen, Z., Liu, Z., Ye, S., Gu, L., Duan, Y., Tian, H., Su, W., Shao, J., et al · 2025
Closest in time.