Fetching the paper…
Reading the bibliography…
Planning is central to agents and agentic AI.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks, 2020a
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 1912
Earlier work this paper cites.
Towards a foundation for evaluating ai planners
Nabil A Kartam and David E Wilkins · 1990
Earlier work this paper cites.
On the complexity of blocks-world planning
Naresh Gupta and Dana S. Nau · 1992
Earlier work this paper cites.
Artificial Intelligence: A modern approach
Stuart Russell and Peter Norvig · 1995
Earlier work this paper cites.
The deterministic part of ipc-4: An overview
J. Hoffmann and S. Edelkamp · 2005
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2010
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning, 2021
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2010
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs, 2018
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba · 2018
Earlier work this paper cites.
Teach: Task-driven embodied agents that chat, 2021
Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gokhan Tur, and Dilek Hakkani-Tur · 2021
Earlier work this paper cites.
Learning to decompose and organize complex tasks
Yi Zhang, Sujay Kumar Jauhar, Julia Kiseleva, Ryen White, and Dan Roth · 2021
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge, 2022
Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar · 2022
Earlier work this paper cites.
Benchmarking the spectrum of agent capabilities, 2022
Danijar Hafner · 2022
Earlier work this paper cites.
What do large language models learn about scripts?
Abhilasha Sancheti and Rachel Rudinger · 2022
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web, 2023
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model, 2023
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Earlier work this paper cites.
Agentbench: Evaluating llms as agents, 2023
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang · 2023
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants, 2023
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom · 2023
Earlier work this paper cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought, 2023
Abulhair Saparov and He He · 2023
Earlier work this paper cites.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Cited alongside, same era.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Cited alongside, same era.
Ruoyao Wang, Graham Todd, Eric Yuan, Ziang Xiao, Marc-Alexandre Côté, and Peter Jansen · 2023
Cited alongside, same era.
Windows agent arena: Evaluating multi-modal os agents at scale, 2024
Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Jang, and Zack Hui · 2024
Cited alongside, same era.
Smartplay: A benchmark for llms as intelligent agents, 2024
Yue Wu, Xuan Tang, Tom M. Mitchell, and Yuanzhi Li · 2024
Later among the works it cites.
Agentgym: Evolving large language model-based agents across diverse environments, 2024
Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Dingwen Yang, Chenyang Liao, Xin Guo, Wei He, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, and Yu-Gang Jiang · 2024
Later among the works it cites.
Theagentcompany: Benchmarking llm agents on consequential real world tasks, 2024
Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Z. Wang, Xuhui Zhou, Zhitong Guo, Murong Cao, Mingyang Yang, Hao Yang Lu, Amaad Martin, Zhe Su, Leander Maben, Raj Mehta, Wayne Chi, Lawrence Jang, Yiqing Xie, Shuyan Zhou, and Graham Neubig · 2024
Later among the works it cites.
Assistantbench: Can web agents solve realistic and time-consuming tasks?, 2024
Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, and Jonathan Berant · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D. Hwang, Xiang Lorraine Li, Hirona J. Arai, Soumya Sanyal, Keisuke Sakaguchi, Xiang Ren, and Yejin Choi · 2024
Cited alongside, same era.
Partnr: A benchmark for planning and reasoning in embodied multi-agent tasks, 2024
Matthew Chang, Gunjan Chhablani, Alexander Clegg, Mikael Dallaire Cote, Ruta Desai, Michal Hlavac, Vladimir Karashchuk, Jacob Krantz, Roozbeh Mottaghi, Priyam Parashar, Siddharth Patki, Ishita Prasad, Xavier Puig, Akshara Rai, Ram Ramrakhya, Daniel Tran, Joanne Truong, John M. Turner, Eric Undersander, and Tsung-Yen Yang · 2024
Cited alongside, same era.
Plancraft: an evaluation dataset for planning with llm agents, 2024
Gautier Dagan, Frank Keller, and Alex Lascarides · 2024
Cited alongside, same era.
Gtbench: Uncovering the strategic reasoning limitations of llms via game-theoretic evaluations, 2024
Jinhao Duan, Renming Zhang, James Diffenderfer, Bhavya Kailkhura, Lichao Sun, Elias Stengel-Eskin, Mohit Bansal, Tianlong Chen, and Kaidi Xu · 2024
Cited alongside, same era.
Game-theoretic llm: Agent workflow for negotiation games, 2024
Wenyue Hua, Ollie Liu, Lingyao Li, Alfonso Amayuelas, Julie Chen, Lucas Jiang, Mingyu Jin, Lizhou Fan, Fei Sun, William Wang, Xintong Wang, and Yongfeng Zhang · 2024
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?, 2024
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2024
Cited alongside, same era.
To the globe (ttg): Towards language-driven guaranteed travel planning, 2024
Da Ju, Song Jiang, Andrew Cohen, Aaron Foss, Sasha Mitts, Arman Zharmagambetov, Brandon Amos, Xian Li, Justine T Kao, Maryam Fazel-Zarandi, and Yuandong Tian · 2024
Cited alongside, same era.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks, 2024
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried · 2024
Cited alongside, same era.
Later among the works it cites.
TimeArena: Shaping efficient multitasking language agents in a time-aware simulation
Yikai Zhang, Siyu Yuan, Caiyu Hu, Kyle Richardson, Yanghua Xiao, and Jiangjie Chen · 2024
Later among the works it cites.
Natural plan: Benchmarking llms on natural language planning, 2024
Huaixiu Steven Zheng, Swaroop Mishra, Hugh Zhang, Xinyun Chen, Minmin Chen, Azade Nova, Le Hou, Heng-Tze Cheng, Quoc V. Le, Ed H. Chi, and Denny Zhou · 2024
Later among the works it cites.
Mohamed Aghzal, Erion Plaku, and Ziyu Yao · 2025
Closest in time.
Workarena++: Towards compositional planning and reasoning-based common knowledge work tasks, 2025
Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault Le Sellier De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin · 2025
Closest in time.
Realm-bench: A real-world planning benchmark for llms and multi-agent systems, 2025
Longling Geng and Edward Y. Chang · 2025
Closest in time.
Robotouille: An asynchronous planning benchmark for llm agents, 2025
Gonzalo Gonzalez-Pumariega, Leong Su Yean, Neha Sunkara, and Sanjiban Choudhury · 2025
Closest in time.
Yilun Hao, Yang Zhang, and Chuchu Fan · 2025
Closest in time.
Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, and Chien-Sheng Wu · 2025
Closest in time.
Large language model strategic reasoning evaluation through behavioral game theory, 2025
Jingru Jia, Zehua Yuan, Junhao Pan, Paul E. McNamara, and Deming Chen · 2025
Closest in time.
Llamar: Long-horizon planning for multi-agent robots in partially observable environments, 2025
Siddharth Nayak, Adelmo Morrison Orozco, Marina Ten Have, Vittal Thirumalai, Jackson Zhang, Darren Chen, Aditya Kapoor, Eric Robinson, Karthik Gopalakrishnan, James Harrison, Brian Ichter, Anuj Mahajan, and Hamsa Balakrishnan · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models, 2025
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, and Xia Hu · 2025
Closest in time.
PlanGenLLMs: A modern survey of llm planning capabilities, 2025
Hui Wei, Zihao Zhang, Shenghua He, Tian Xia, Shijia Pan, and Fei Liu · 2025
Closest in time.
Zirui Wu, Xiao Liu, Jiayi Li, Lingpeng Kong, and Yansong Feng · 2025
Closest in time.
Safeagentbench: A benchmark for safe task planning of embodied llm agents, 2025
Sheng Yin, Xianghe Pang, Yuanzhuo Ding, Menglan Chen, Yutong Bi, Yichen Xiong, Wenhao Huang, Zhen Xiang, Jing Shao, and Siheng Chen · 2025
Closest in time.