Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) have been applied to robot task planning problems, where the robot receives a task in natural language and generates plans based on visual inputs.
Shakey the robot
N. J. Nilsson et al · 1984
Earlier work this paper cites.
Answer set programming and plan generation
V. Lifschitz · 2002
Earlier work this paper cites.
Pddl2. 1: An extension to pddl for expressing temporal planning domains
M. Fox and D. Long · 2003
Earlier work this paper cites.
D. Driess, J.-S. Ha, and M. Toussaint · 2006
Earlier work this paper cites.
The fast downward planning system
M. Helmert · 2006
Earlier work this paper cites.
Answer set programming at a glance
G. Brewka, T. Eiter, and M. Truszczyński · 2011
Earlier work this paper cites.
Integrated task and motion planning in belief space
L. P. Kaelbling and T. Lozano-Pérez · 2013
Earlier work this paper cites.
Mobile robot planning using action language bc with an abstraction hierarchy
S. Zhang, F. Yang, P. Khandelwal, and P. Stone · 2015
Earlier work this paper cites.
Automated planning and acting
M. Ghallab, D. Nau, and P. Traverso · 2016
Earlier work this paper cites.
Platform-independent benchmarks for task and motion planning
F. Lagriffoul, N. T. Dantam, C. Garrett, A. Akbari, S. Srivastava, and L. E. Kavraki · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Closing the loop for robotic grasping: A real-time, generative grasp synthesis approach
D. Morrison, P. Corke, and J. Leitner · 2018
Earlier work this paper cites.
Task planning in robotics: an empirical comparison of pddl-and asp-based systems
Y.-q. Jiang, S.-q. Zhang, P. Khandelwal, and P. Stone · 2019
Earlier work this paper cites.
Multi-robot planning with conflicts and synergies
Y. Jiang, H. Yedidsion, S. Zhang, G. Sharon, and P. Stone · 2019
Earlier work this paper cites.
Learning feasibility for task and motion planning in tabletop environments
A. M. Wells, N. T. Dantam, A. Shrivastava, and L. E. Kavraki · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Task-motion planning for safe and efficient urban driving
Y. Ding, X. Zhang, X. Zhan, and S. Zhang · 2020
Earlier work this paper cites.
Deep visual heuristics: Learning feasibility of mixed-integer programs for manipulation planning
D. Driess, O. Oguz, J.-S. Ha, and M. Toussaint · 2020
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Earlier work this paper cites.
Hierarchical planning for long-horizon manipulation with geometric and symbolic scene graphs
Y. Zhu, J. Tremblay, S. Birchfield, and Y. Zhu · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Learning to ground objects for robot task and motion planning
Y. Ding, X. Zhang, X. Zhan, and S. Zhang · 2022
Earlier work this paper cites.
Visually grounded task and motion planning for mobile manipulation
X. Zhang, Y. Zhu, Y. Ding, Y. Zhu, P. Stone, and S. Zhang · 2022
Cited alongside, same era.
Grounding predicates through actions
T. Migimatsu and J. Bohg · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Cited alongside, same era.
Housekeep: Tidying virtual households using commonsense reasoning
Y. Kant, A. Ramachandran, S. Yenamandra, I. Gilitschenski, D. Batra, A. Szot, and H. Agrawal · 2022
Cited alongside, same era.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Cited alongside, same era.
Llm+ p: Empowering large language models with optimal planning proficiency
B. Liu, Y. Jiang, X. Zhang, Q. Liu, S. Zhang, J. Biswas, and P. Stone · 2023
Later among the works it cites.
Autoplanbench:: Automatically generating benchmarks for llm planners from pddl
K. Stein and A. Koller · 2023
Later among the works it cites.
Leveraging pre-trained large language models to construct and utilize world models for model-based task planning
L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati · 2023
Later among the works it cites.
Task and motion planning with large language models for object rearrangement
Y. Ding, X. Zhang, C. Paxton, and S. Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, et al · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al · 2022
Cited alongside, same era.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2022
Cited alongside, same era.
Structdiffusion: Object-centric diffusion for semantic rearrangement of novel objects
W. Liu, T. Hermans, S. Chernova, and C. Paxton · 2022
Cited alongside, same era.
Large language models still can’t plan (a benchmark for llms on planning and reasoning about change)
K. Valmeekam, A. Olmo, S. Sreedharan, and S. Kambhampati · 2022
Cited alongside, same era.
Pddl planning with pretrained large language models
T. Silver, V. Hariprasad, R. S. Shuttleworth, N. Kumar, T. Lozano-Pérez, and L. P. Kaelbling · 2022
Cited alongside, same era.
Plansformer: Generating symbolic plans using transformers
V. Pallagani, B. Muppasani, K. Murugesan, F. Rossi, L. Horesh, B. Srivastava, F. Fabiano, and A. Loreggia · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Gpt-4v (ision) for robotics: Multimodal task planning from human demonstration
N. Wake, A. Kanehira, K. Sasabuchi, J. Takamatsu, and K. Ikeuchi · 2023
Later among the works it cites.
Robovqa: Multimodal long-horizon reasoning for robotics
P. Sermanet, T. Ding, J. Zhao, F. Xia, D. Dwibedi, K. Gopalakrishnan, C. Chan, G. Dulac-Arnold, S. Maddineni, N. J. Joshi, et al · 2023
Later among the works it cites.
Open-world object manipulation using pre-trained vision-language model
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, B. Zitkovich, F. Xia, C. Finn, and K. Hausman · 2023
Later among the works it cites.
Chat with the environment: Interactive multimodal perception using large language models
X. Zhao, M. Li, C. Weber, M. B. Hafez, and S. Wermter · 2023
Later among the works it cites.
Robots that ask for help: Uncertainty alignment for large language model planners
A. Z. Ren, A. Dixit, A. Bodrova, S. Singh, S. Tu, N. Brown, P. Xu, L. Takayama, F. Xia, J. Varley, et al · 2023
Later among the works it cites.
Vision-language models as success detectors
Y. Du, K. Konyushkova, M. Denil, A. Raju, J. Landon, F. Hill, N. de Freitas, and S. Cabi · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Later among the works it cites.
Doremi: Grounding language model by detecting and recovering from plan-execution misalignment
Y. Guo, Y.-J. Wang, L. Zha, Z. Jiang, and J. Chen · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Closest in time.
Generalized planning in pddl domains with pretrained large language models
T. Silver, S. Dan, K. Srinivas, J. B. Tenenbaum, L. Kaelbling, and M. Katz · 2024
Closest in time.
Llmˆ 3: Large language model-based task and motion planning with motion failure reasoning
S. Wang, M. Han, Z. Jiao, Z. Zhang, Y. N. Wu, S.-C. Zhu, and H. Liu · 2024
Closest in time.
Mm-llms: Recent advances in multimodal large language models
D. Zhang, Y. Yu, C. Li, J. Dong, D. Su, C. Chu, and D. Yu · 2024
Closest in time.
Claude 3 family, 2023
Anthropic · 2024
Closest in time.
Cognitivedog: Large multimodal model based system to translate vision and language into action of quadruped robot
A. Lykov, M. Litvinov, M. Konenkov, R. Prochii, N. Burtsev, A. A. Abdulkarim, A. Bazhenov, V. Berman, and D. Tsetserukou · 2024
Closest in time.
L. Guan, Y. Zhou, D. Liu, Y. Zha, H. B. Amor, and S. Kambhampati · 2024
Closest in time.
Openeqa: Embodied question answering in the era of foundation models
A. Majumdar, A. Ajay, X. Zhang, P. Putta, S. Yenamandra, M. Henaff, S. Silwal, P. Mcvay, O. Maksymets, S. Arnaud, et al · 2024
Closest in time.
Robomp 2 2 : A robotic multimodal perception-planning framework with mutlimodal large language models
Q. Lv, H. Li, X. Deng, R. Shao, M. Y. Wang, and L. Nie · 2024
Closest in time.
Closed-loop open-vocabulary mobile manipulation with gpt-4v, 2024
P. Zhi, Z. Zhang, M. Han, Z. Zhang, Z. Li, Z. Jiao, B. Jia, and S. Huang · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.