Fetching the paper…
Reading the bibliography…
Recent advances in large vision-language models (LVLMs) have shown promise for embodied task planning, yet they struggle with fundamental challenges like dependency constraints and efficiency.
Cognitive maps in rats and men
Edward C Tolman · 1948
Earlier work this paper cites.
Mental models: Towards a cognitive science of language, inference, and consciousness
Philip Nicholas Johnson-Laird · 1983
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S. Sutton · 1990
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, Aniruddha Kembhavi, Abhinav Kumar Gupta, and Ali Farhadi · 2017
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy P. Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2019
Earlier work this paper cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2019
Earlier work this paper cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Earlier work this paper cites.
Episodic transformer for vision-and-language navigation
Alexander Pashevich, Cordelia Schmid, and Chen Sun · 2021
Earlier work this paper cites.
Yuki Inoue and Hiroki Ohashi · 2022
Earlier work this paper cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Yann LeCun · 2022
Earlier work this paper cites.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, F. Xia, Peng Xu, Karol Hausman, Brian Ichter, Peter R. Florence, and Andy Zeng · 2022
Earlier work this paper cites.
Progprompt: Generating situated robot task plans using large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg · 2022
Earlier work this paper cites.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Jiaman Wu, Clay Washington, Brian M. Sadler, Wei-Lun Chao, and Yu Su · 2022
Earlier work this paper cites.
Daydreamer: World models for physical robot learning
Philipp Wu, Alejandro Escontrela, Danijar Hafner, Ken Goldberg, and P. Abbeel · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2022
Earlier work this paper cites.
Grounding large language models in interactive environments with online reinforcement learning
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2023
Earlier work this paper cites.
L. Guan, Karthik Valmeekam, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Cited alongside, same era.
Mastering diverse domains through world models
Danijar Hafner, J. Pasukonis, Jimmy Ba, and Timothy P. Lillicrap · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Cited alongside, same era.
Alphablock: Embodied finetuning for vision-language reasoning in robot manipulation
Chuhao Jin, Wenhui Tan, Jiange Yang, Bei Liu, Ruihua Song, Limin Wang, and Jianlong Fu · 2023
Cited alongside, same era.
Generating code world models with large language models guided by monte carlo tree search
Nicola Dainese, Matteo Merler, Minttu Alakuijala, and Pekka Marttinen · 2024
Later among the works it cites.
Selu: Self-learning embodied mllms in unknown environments
Boyu Li, Haobin Jiang, Ziluo Ding, Xinrun Xu, Haoran Li, Dongbin Zhao, and Zongqing Lu · 2024
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Later among the works it cites.
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
AI Meta · 2024
Later among the works it cites.
OpenAI · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guanxing Lu, Ziwei Wang, Changliu Liu, Jiwen Lu, and Yansong Tang · 2023
Cited alongside, same era.
Llm as a robotic brain: Unifying egocentric memory and control
Jinjie Mai, Jun Chen, Bing chuan Li, Guocheng Qian, Mohamed Elhoseiny, and Bernard Ghanem · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Cited alongside, same era.
Vision-language interpreter for robot task planning
Keisuke Shirai, Cristian Camilo Beltran-Hernandez, Masashi Hamaya, Atsushi Hashimoto, Shohei Tanaka, Kento Kawaharazuka, Kazutoshi Tanaka, Yoshitaka Ushiku, and Shinsuke Mori · 2023
Cited alongside, same era.
Adaplanner: Adaptive planning from feedback with language models
Haotian Sun, Yuchen Zhuang, Lingkai Kong, Bo Dai, and Chao Zhang · 2023
Cited alongside, same era.
Large language models as generalizable policies for embodied tasks
Andrew Szot, Max Schwarzer, Harsh Agrawal, Bogdan Mazoure, Walter Talbott, Katherine Metcalf, Natalie Mackraz, Devon Hjelm, and Alexander Toshev · 2023
Cited alongside, same era.
Embodied task planning with large language models
Zhenyu Wu, Ziwei Wang, Xiuwei Xu, Jiwen Lu, and Haibin Yan · 2023
Cited alongside, same era.
Octopus: Embodied vision-language programmer from environmental feedback
Jingkang Yang, Yuhao Dong, Shuai Liu, Bo Li, Ziyue Wang, Chencheng Jiang, Haoran Tan, Jiamu Kang, Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu · 2023
Cited alongside, same era.
Later among the works it cites.
Socratic planner: Inquiry-based zero-shot planning for embodied instruction following
Suyeon Shin, Sujin Jeon, Junghyun Kim, Gi-Cheon Kang, and Byoung-Tak Zhang · 2024
Later among the works it cites.
Trial and error: Exploration-based trajectory optimization for llm agents
Yifan Song, Da Yin, Xiang Yue, Jie Huang, Sujian Li, and Bill Yuchen Lin · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team · 2024
Later among the works it cites.
V-dpo: Mitigating hallucination in large vision language models via vision-guided direct preference optimization
Yuxi Xie, Guanzhen Li, Xiao Xu, and Min-Yen Kan · 2024
Later among the works it cites.
Hindsight planner: A closed-loop few-shot planner for embodied instruction following
Yuxiao Yang, Shenao Zhang, Zhihan Liu, Huaxiu Yao, and Zhaoran Wang · 2024
Later among the works it cites.
Epo: Hierarchical llm agents with environment preference optimization
Qi Zhao, Haotian Fu, Chen Sun, and George Dimitri Konidaris · 2024
Later among the works it cites.
Wall-e: World alignment by rule learning improves world model-based llm agents
Siyu Zhou, Tianyi Zhou, Yijun Yang, Guodong Long, Deheng Ye, Jing Jiang, and Chengqi Zhang · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Chip: Cross-modal hierarchical direct preference optimization for multimodal llms
Jinlan Fu, Shenzhen Huangfu, Hao Fei, Xiaoyu Shen, Bryan Hooi, Xipeng Qiu, and See-Kiong Ng · 2025
Closest in time.
Videoworld: Exploring knowledge learning from unlabeled videos, 2025
Zhongwei Ren, Yunchao Wei, Xun Guo, Yao Zhao, Bingyi Kang, Jiashi Feng, and Xiaojie Jin · 2025
Closest in time.