Fetching the paper…
Reading the bibliography…
The capability of Large Language Models (LLMs) to plan remains a topic of debate.
Go-explore: a new approach for hard-exploration problems
Ecoffet, A.; Huizinga, J.; Lehman, J.; Stanley, K. O.; and Clune, J. 2019 · 1901
Earlier work this paper cites.
Randaugment: Practical data augmentation with no separate search
Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2019 · 1909
Earlier work this paper cites.
On the NP-Hardness of Blocks World
Chenoweth, S. V. 1991 · 1991
Earlier work this paper cites.
Affinity and diversity: Quantifying mechanisms of data augmentation
Gontijo-Lopes, R.; Smullin, S. J.; Cubuk, E. D.; and Dyer, E. 2020 · 2002
Earlier work this paper cites.
VAL: Automatic Plan Validation, Continuous Effects and Mixed Initiative Planning Using PDDL
Howey, R.; et al. 2004 · 2004
Earlier work this paper cites.
Chen, G.; Ding, Y.; Edwards, H.; Chau, C. H.; Hou, S.; Johnson, G.; Sharukh Syed, M.; Tang, H.; Wu, Y.; Yan, Y.; Gil, T.; and Nir, L. 2020 · 2008
Earlier work this paper cites.
An introduction to the planning domain definition language , volume 13
Haslum, P.; Lipovetzky, N.; Magazzeni, D.; Muise, C.; Brachman, R.; Rossi, F.; and Stone, P. 2019 · 2019
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L.; Lu, K.; Rajeswaran, A.; Lee, K.; Grover, A.; Laskin, M.; Abbeel, P.; Srinivas, A.; and Mordatch, I. 2021 · 2021
Earlier work this paper cites.
Cont: Contrastive neural text generation
An, C.; Feng, J.; Lv, K.; Kong, L.; Qiu, X.; and Huang, X. 2022 · 2022
Earlier work this paper cites.
PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
Valmeekam, K.; Olmo, A.; Sreedharan, S.; and Kambhampati, S. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Earlier work this paper cites.
Physics of language models: Part 3.1, knowledge storage and extraction
Allen-Zhu, Z.; and Li, Y. 2023 · 2023
Earlier work this paper cites.
Huang, L.; Yu, W.; Ma, W.; Zhong, W.; Feng, Z.; Wang, H.; Chen, Q.; Peng, W.; Feng, X.; Qin, B.; and Liu, T. 2023 · 2023
Cited alongside, same era.
Tinygsm: achieving¿ 80% on gsm8k with small language models
Liu, B.; Bubeck, S.; Eldan, R.; Kulkarni, J.; Li, Y.; Nguyen, A.; Ward, R.; and Zhang, Y. 2023 · 2023
Cited alongside, same era.
Why think step by step? Reasoning emerges from the locality of experience
Prystawski, B.; Li, M.; and Goodman, N. D. 2023 · 2023
Cited alongside, same era.
Language models are multilingual chain-of-thought reasoners
Shi, F.; Suzgun, M.; Freitag, M.; Wang, X.; Srivats, S.; Vosoughi, S.; Chung, H. W.; Tay, Y.; Ruder, S.; Zhou, D.; Das, D.; and Wei, J. 2023 · 2023
Cited alongside, same era.
On the planning abilities of large language models-a critical investigation
Valmeekam, K.; Marquez, M.; Sreedharan, S.; and Kambhampati, S. 2023 · 2023
Thought of Search: Planning with Language Models Through The Lens of Efficiency
Katz, M.; Kokel, H.; Srinivas, K.; and Sohrabi, S. 2024 · 2024
Closest in time.
Training Language Models to Self-Correct via Reinforcement Learning
Kumar, A.; Zhuang, V.; Agarwal, R.; Su, Y.; Co-Reyes, J. D.; Singh, A.; Baumli, K.; Iqbal, S.; Bishop, C.; Roelofs, R.; et al. 2024 · 2024
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
Liu, A.; Feng, B.; Wang, B.; Wang, B.; Liu, B.; Zhao, C.; Dengr, C.; Ruan, C.; Dai, D.; Guo, D.; et al. 2024 · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. 2024 · 2024
Closest in time.
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Ahmadian, A.; Cremer, C.; Gallé, M.; Fadaee, M.; Kreutzer, J.; Üstün, A.; and Hooker, S. 2024 · 2024
Cited alongside, same era.
The pitfalls of next-token prediction
Bachmann, G.; et al. 2024 · 2024
Cited alongside, same era.
Automating Thought of Search: A Journey Towards Soundness and Completeness
Cao, D.; Katz, M.; Kokel, H.; Srinivas, K.; and Sohrabi, S. 2024 · 2024
Cited alongside, same era.
AlphaMath Almost Zero: Process Supervision without Process
Chen, G.; Liao, M.; Li, C.; and Fan, K. 2024 · 2024
Cited alongside, same era.
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024 · 2024
Cited alongside, same era.
Learning Mathematical Rules with Large Language Models
Gorceix, A.; Le Chenadec, B.; Rammal, A.; Vadori, N.; and Veloso, M. 2024 · 2024
Cited alongside, same era.
Position: LLMs Can’t Plan, But Can Help Planning in LLM-Modulo Frameworks
Kambhampati, S.; Valmeekam, K.; Guan, L.; Verma, M.; Stechly, K.; Bhambri, S.; Saldyt, L. P.; and Murthy, A. B. 2024 · 2024
Cited alongside, same era.
Mirzadeh, I.; Alizadeh, K.; Shahrokhi, H.; Tuzel, O.; Bengio, S.; and Farajtabar, M. 2024 · 2024
Closest in time.
Learning General Policies for Planning through GPT Models
Rossetti, N.; Tummolo, M.; Gerevini, A. E.; Putelli, L.; Serina, I.; Chiari, M.; and Olivato, M. 2024 · 2024
Closest in time.
Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles
Shah, K.; Dikkala, N.; Wang, X.; and Panigrahy, R. 2024 · 2024
Closest in time.
Chain of thoughtlessness: An analysis of CoT in planning
Stechly, K.; et al. 2024 · 2024
Closest in time.
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; et al. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; and Narasimhan, K. 2024 · 2024
Closest in time.
Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems
Ye, T.; Xu, Z.; Li, Y.; and Allen-Zhu, Z. 2024 · 2024
Closest in time.