Fetching the paper…
Reading the bibliography…
Advancing planning and reasoning capabilities of Large Language Models (LLMs) is one of the key prerequisites towards unlocking their potential for performing reliably in complex and impactful domains.
Scientific American Supplement , 80(2079):296, November 1915
Torres and his remarkable automatic devices — he would substitute machinery for the human mind · 1915
Earlier work this paper cites.
The proposed uscf rating system, its development, theory, and applications
A. E. Elo · 1967
Earlier work this paper cites.
Chess as the drosophila of ai
J. McCarthy · 1990
Earlier work this paper cites.
Parallel Monte-Carlo tree search
G. M. B. Chaslot, M. H. Winands, and H. J. van Den Herik · 2008
Earlier work this paper cites.
Progressive strategies for monte carlo tree search
G.-B. Chaslot, M. Winands, H. van den Herik, J. Uiterwijk, and B. Bouzy · 2009
Earlier work this paper cites.
Thinking, Fast and Slow
D. Kahneman · 2011
Earlier work this paper cites.
Measuring intelligence through games
T. Schaul, J. Togelius, and J. Schmidhuber · 2011
Earlier work this paper cites.
Is chess the drosophila of artificial intelligence? a social history of an algorithm
N. Ensmenger · 2012
Earlier work this paper cites.
An analysis of virtual loss in parallel MCTS
S. A. Mirsoleimani, A. Plaat, J. van den Herik, and J. Vermaseren · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis · 2017
Earlier work this paper cites.
Openspiel: A framework for reinforcement learning in games
M. Lanctot, E. Lockhart, J.-B. Lespiau, V. Zambaldi, S. Upadhyay, J. Pérolat, S. Srinivasan, F. Timbers, K. Tuyls, S. Omidshafiei, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
T. Gao, A. Fisch, and D. Chen · 2020
Earlier work this paper cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. Lillicrap, and D. Silver · 2020
Earlier work this paper cites.
Geometric deep learning: Grids, groups, graphs, geodesics, and gauges
M. M. Bronstein, J. Bruna, T. Cohen, and P. Veličković · 2021
Earlier work this paper cites.
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Earlier work this paper cites.
In-context reinforcement learning with algorithm distillation, 2022
M. Laskin, L. Wang, J. Oh, E. Parisotto, S. Spencer, R. Steigerwald, D. Strouse, S. Hansen, A. Filos, E. Brooks, M. Gazeau, H. Sahni, S. Singh, and V. Mnih · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Earlier work this paper cites.
Measuring and narrowing the compositionality gap in language models
O. Press, M. Zhang, S. Min, L. Schmidt, N. A. Smith, and M. Lewis · 2022
Earlier work this paper cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
A. Saparov and H. He · 2022
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners
F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, et al · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Earlier work this paper cites.
The unreliability of explanations in few-shot prompting for textual reasoning
X. Ye and G. Durrett · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning, 2022
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman · 2022
Earlier work this paper cites.
Automatic chain of thought prompting in large language models, 2022
Z. Zhang, A. Zhang, M. Li, and A. Smola · 2022
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. Le, et al · 2022
Earlier work this paper cites.
Debunking the chessboard: Confronting gpts against chess engines to estimate elo ratings and assess legal move abilities, 2023
M. Acher · 2023
Earlier work this paper cites.
Rest meets react: Self-improvement for multi-step reasoning llm agent, 2023
R. Aksitov, S. Miryoosefi, Z. Li, D. Li, S. Babayan, K. Kopparapu, Z. Fisher, R. Guo, S. Prakash, P. Srinivasan, M. Zaheer, F. Yu, and S. Kumar · 2023
Earlier work this paper cites.
Y. Bang, S. Cahyawijaya, N. Lee, W. Dai, D. Su, B. Wilie, H. Lovenia, Z. Ji, T. Yu, W. Chung, Q. V. Do, Y. Xu, and P. Fung · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y. Chen, A. Ecoffet, M. Joglekar, J. Leike, et al · 2023
Earlier work this paper cites.
Dynamic planning with a llm, 2023
G. Dagan, F. Keller, and A. Lascarides · 2023
Earlier work this paper cites.
Everything of thoughts: Defying the law of penrose triangle for thought generation
R. Ding, C. Zhang, L. Wang, Y. Xu, M. Ma, W. Zhang, S. Qin, S. Rajmohan, Q. Lin, and D. Zhang · 2023
Earlier work this paper cites.
Chessgpt: Bridging policy learning and language modeling, 2023
X. Feng, Y. Luo, Z. Wang, H. Tang, M. Yang, K. Shao, D. Mguni, Y. Du, and J. Wang · 2023
Earlier work this paper cites.
Strategic reasoning with language models
K. Gandhi, D. Sadigh, and N. D. Goodman · 2023
Earlier work this paper cites.
A survey on large language models: Applications, challenges, limitations, and practical usage
M. U. Hadi, R. Qureshi, A. Shah, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wu, S. Mirjalili, et al · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu · 2023
Earlier work this paper cites.
Large language models as simulated economic agents: What can we learn from homo silicus?
J. J. Horton · 2023
Cited alongside, same era.
Large language model-empowered agents for simulating macroeconomic activities
N. Li, C. Gao, Y. Li, and Q. Liao · 2023
Cited alongside, same era.
Large language model guided tree-of-thought
J. Long · 2023
Cited alongside, same era.
Bridging chess mastery and ai innovation: The making of llm-chesscoach, 2023
A. Menon · 2023
Cited alongside, same era.
Art: Automatic multi-step reasoning and tool-use for large language models
B. Paranjape, S. Lundberg, S. Singh, H. Hajishirzi, L. Zettlemoyer, and M. T. Ribeiro · 2023
Cited alongside, same era.
Training language models to self-correct via reinforcement learning
A. Kumar, V. Zhuang, R. Agarwal, Y. Su, J. D. Co-Reyes, A. Singh, K. Baumli, S. Iqbal, C. Bishop, R. Roelofs, et al · 2024
Closest in time.
Can small language models help large language models reason better?: Lm-guided chain-of-thought
J. Lee, F. Yang, T. Tran, Q. Hu, E. Barut, K.-W. Chang, and C. Su · 2024
Closest in time.
Beyond a ∗ : Better planning with transformers via search dynamics bootstrapping, 2024
L. Lehnert, S. Sukhbaatar, D. Su, Q. Zheng, P. Mcvay, M. Rabbat, and Y. Tian · 2024
Closest in time.
Strategist: Learning strategic skills by llms via bi-level tree search, 2024
J. Light, M. Cai, W. Chen, G. Wang, X. Chen, W. Cheng, Y. Yue, and Z. Hu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Synthetic prompting: Generating chain-of-thought demonstrations for large language models
Z. Shao, Y. Gong, Y. Shen, M. Huang, N. Duan, and W. Chen · 2023
Cited alongside, same era.
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
L. Wang, W. Xu, Y. Lan, Z. Hu, Y. Lan, R. K.-W. Lee, and E.-P. Lim · 2023
Cited alongside, same era.
Multimodal chain-of-thought reasoning in language models
Z. Zhang, A. Zhang, M. Li, H. Zhao, G. Karypis, and A. Smola · 2023
Cited alongside, same era.
R. Agarwal, A. Singh, L. M. Zhang, B. Bohnet, S. Chan, A. Anand, Z. Abbas, A. Nova, J. D. Co-Reyes, E. Chu, et al · 2024
Cited alongside, same era.
Implicit search via discrete diffusion: A study on chess
Anonymous · 2024
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk, and T. Hoefler · 2024
Cited alongside, same era.
Reliable reasoning beyond natural language, 2024
N. Borazjanizadeh and S. T. Piantadosi · 2024
Cited alongside, same era.
P. Lu, B. Peng, H. Cheng, M. Galley, K.-W. Chang, Y. N. Wu, S.-C. Zhu, and J. Gao · 2024
Closest in time.
The clrs-text algorithmic reasoning language benchmark
L. Markeeva, S. McLeish, B. Ibarz, W. Bounsi, O. Kozlova, A. Vitvitskyi, C. Blundell, T. Goldstein, A. Schwarzschild, and P. Veličković · 2024
Closest in time.
E. Markowitz, A. Ramakrishna, J. Dhamala, N. Mehrabi, C. Peris, R. Gupta, K.-W. Chang, and A. Galstyan · 2024
Closest in time.
Whiteboard-of-thought: Thinking step-by-step across modalities, 2024
S. Menon, R. Zemel, and C. Vondrick · 2024
Closest in time.
Large language models: A survey
S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao · 2024
Closest in time.
Compositional chain-of-thought prompting for large multimodal models
C. Mitra, B. Huang, T. Darrell, and R. Herzig · 2024
Closest in time.
Mastering chess with a transformer model, 2024
D. Monroe and T. Leela Chess Zero Team · 2024
Closest in time.
One step at a time: Language agents are stepwise planners, 2024
M. Nguyen and E. Shareghi · 2024
Closest in time.
H. Nori, N. Usuyama, N. King, S. M. McKinney, X. Fernandes, S. Zhang, and E. Horvitz · 2024
Closest in time.
Dynamic strategy planning for efficient question answering with large language models, 2024
T. Parekh, P. Prakash, A. Radovic, A. Shekher, and D. Savenkov · 2024
Closest in time.
Let’s think dot by dot: Hidden computation in transformer language models, 2024
J. Pfau, W. Merrill, and S. R. Bowman · 2024
Closest in time.
Reasoning with large language models, a survey
A. Plaat, A. Wong, S. Verberne, J. Broekens, N. van Stein, and T. Back · 2024
Closest in time.
Why think step by step? reasoning emerges from the locality of experience
B. Prystawski, M. Li, and N. Goodman · 2024
Closest in time.
Agent q: Advanced reasoning and learning for autonomous ai agents, 2024
P. Putta, E. Mills, N. Garg, S. Motwani, C. Finn, D. Garg, and R. Rafailov · 2024
Closest in time.
From r r to q ∗ q^{*} : Your language model is secretly a q-function, 2024
R. Rafailov, J. Hejna, R. Park, and C. Finn · 2024
Closest in time.
Optimal decision making through scenario simulations using large language models, 2024
S. Rasal and E. J. Hauer · 2024
Closest in time.
D. Rebstock, C. Solinas, N. R. Sturtevant, and M. Buro · 2024
Closest in time.
Thinking forward and backward: Effective backward planning with large language models, 2024
A. Z. Ren, B. Ichter, and A. Majumdar · 2024
Closest in time.
Capabilities of gemini models in medicine
K. Saab, T. Tu, W.-H. Weng, R. Tanno, D. Stutz, E. Wulczyn, F. Zhang, T. Strother, C. Park, E. Vedadi, et al · 2024
Closest in time.
Math-llava: Bootstrapping mathematical reasoning for multimodal large language models, 2024
W. Shi, Z. Hu, Y. Bin, J. Liu, Y. Yang, S.-K. Ng, L. Bing, and R. K.-W. Lee · 2024
Closest in time.
Beyond human data: Scaling self-training for problem-solving with language models, 2024
A. Singh, J. D. Co-Reyes, R. Agarwal, A. Anand, P. Patil, X. Garcia, P. J. Liu, J. Harrison, J. Lee, K. Xu, A. Parisi, A. Kumar, A. Alemi, A. Rizkowsky, A. Nova, B. Adlam, B. Bohnet, G. Elsayed, H. Sedghi, I. Mordatch, I. Simpson, I. Gur, J. Snoek, J. Pennington, J. Hron, K. Kenealy, K. Swersky, K. Mahajan, L. Culp, L. Xiao, M. L. Bileschi, N. Constant, R. Novak, R. Liu, T. Warkentin, Y. Qian, Y. Bansal, E. Dyer, B. Neyshabur, J. Sohl-Dickstein, and N. Fiedel · 2024
Closest in time.
Chain of thoughtlessness? an analysis of cot in planning, 2024
K. Stechly, K. Valmeekam, and S. Kambhampati · 2024
Closest in time.
Toward self-improvement of llms via imagination, searching, and criticizing, 2024
Y. Tian, B. Peng, L. Song, L. Jin, D. Yu, H. Mi, and D. Yu · 2024
Closest in time.
Benchmarking large language model (llm) performance for game playing via tic-tac-toe
O. Topsakal and J. B. Harper · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
M. Turpin, J. Michael, E. Perez, and S. Bowman · 2024
Closest in time.
softmax is not enough (for sharp out-of-distribution)
P. Veličković, C. Perivolaropoulos, F. Barbero, and R. Pascanu · 2024
Closest in time.
Large language models can learn temporal reasoning
S. Xiong, A. Payani, R. Kompella, and F. Fekri · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Closest in time.
Quiet-star: Language models can teach themselves to think before speaking, 2024
E. Zelikman, G. Harik, Y. Shao, V. Jayasiri, N. Haber, and N. D. Goodman · 2024
Closest in time.
MR-ben: A meta-reasoning benchmark for evaluating system-2 thinking in LLMs
Z. Zeng, Y. Liu, Y. Wan, J. Li, P. Chen, J. Dai, Y. Yao, R. Xu, Z. Qi, W. Zhao, L. Shen, J. Lu, H. Tan, Y. Chen, H. Zhang, Z. Shi, B. Wang, Z. Guo, and J. Jia · 2024
Closest in time.
Large language models as commonsense knowledge for large-scale task planning
Z. Zhao, W. S. Lee, and D. Hsu · 2024
Closest in time.
Language agent tree search unifies reasoning acting and planning in language models, 2024
A. Zhou, K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y.-X. Wang · 2024
Closest in time.
Can large language models transform computational social science?
C. Ziems, W. Held, O. Shaikh, J. Chen, Z. Zhang, and D. Yang · 2024
Closest in time.