Fetching the paper…
Reading the bibliography…
Enabling LLMs to improve their outputs by using more test-time computation is a critical step towards building generally self-improving agents that can operate on open-ended natural language.
Heuristic and analytic processes in reasoning
J. S. B. T. Evans · 1984
Earlier work this paper cites.
An introduction to mcmc for machine learning
C. Andrieu, N. De Freitas, A. Doucet, and M. I. Jordan · 2003
Earlier work this paper cites.
Maps of bounded rationality: Psychology for behavioral economics
D. Kahneman · 2003
Earlier work this paper cites.
Bandit based monte-carlo planning
L. Kocsis and C. Szepesv’ari · 2006
Earlier work this paper cites.
Thinking, fast and slow
D. Kahneman · 2013
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Scaling scaling laws with board games, 2021
A. L. Jones · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback, 2022
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan · 2022
Earlier work this paper cites.
Training compute-optimal large language models, 2022
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models, 2022
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators, 2022
W. Saunders, C. Yeh, J. Wu, S. Bills, L. Ouyang, J. Ward, and J. Leike · 2022
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback, 2022
J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning, 2022
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman · 2022
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate, 2023
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch · 2023
Earlier work this paper cites.
Pal: Program-aided language models, 2023
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2023
Cited alongside, same era.
Large language models cannot self-correct reasoning yet, 2023
J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou · 2023
Cited alongside, same era.
Making large language models better reasoners with step-aware verifier, 2023
Y. Li, Z. Lin, S. Zhang, Q. Fu, B. Chen, J.-G. Lou, and W. Chen · 2023
Cited alongside, same era.
Let’s verify step by step, 2023
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback, 2023
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark · 2023
Cited alongside, same era.
Does your data spark joy? performance gains from domain upsampling at the end of training, 2024
C. Blakeney, M. Paul, B. W. Larsen, S. Owen, and J. Frankle · 2024
Closest in time.
Alphamath almost zero: process supervision without process, 2024
G. Chen, M. Liao, C. Li, and K. Fan · 2024
Closest in time.
Alphazero-like tree-search can guide large language model decoding and training, 2024
X. Feng, Z. Wan, M. Wen, S. M. McAleer, Y. Wen, W. Zhang, and J. Wang · 2024
Closest in time.
Think before you speak: Training language models with pause tokens, 2024
S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V. Nagarajan · 2024
Closest in time.
Llm critics help catch llm bugs
N. McAleese, R. Pokorny, J. F. Cerón Uribe, E. Nitishinskaya, M. Trębacz, and J. Leike · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, L. Hong, R. Tian, R. Xie, J. Zhou, M. Gerstein, D. Li, Z. Liu, and M. Sun · 2023
Cited alongside, same era.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws, 2023
N. Sardana and J. Frankle · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning, 2023
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Cited alongside, same era.
Gpt-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems, 2023
K. Stechly, M. Marquez, and S. Kambhampati · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Cited alongside, same era.
Can large language models really improve by self-critiquing their own plans?, 2023
K. Valmeekam, M. Marquez, and S. Kambhampati · 2023
Cited alongside, same era.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations, 2023
P. Wang, L. Li, Z. Shao, R. X. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui · 2023
Cited alongside, same era.
Gpt-4 technical report, 2024
OpenAI · 2024
Closest in time.
Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold
A. Setlur, S. Garg, X. Geng, N. Garg, V. Smith, and A. Kumar · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo · 2024
Closest in time.
A critical evaluation of ai feedback for aligning large language models, 2024
A. Sharma, S. Keh, E. Mitchell, C. Finn, K. Arora, and T. Kollar · 2024
Closest in time.
Beyond human data: Scaling self-training for problem-solving with language models, 2024
A. Singh, J. D. Co-Reyes, R. Agarwal, A. Anand, P. Patil, X. Garcia, P. J. Liu, J. Harrison, J. Lee, K. Xu, A. Parisi, A. Kumar, A. Alemi, A. Rizkowsky, A. Nova, B. Adlam, B. Bohnet, G. Elsayed, H. Sedghi, I. Mordatch, I. Simpson, I. Gur, J. Snoek, J. Pennington, J. Hron, K. Kenealy, K. Swersky, K. Mahajan, L. Culp, L. Xiao, M. L. Bileschi, N. Constant, R. Novak, R. Liu, T. Warkentin, Y. Qian, Y. Bansal, E. Dyer, B. Neyshabur, J. Sohl-Dickstein, and N. Fiedel · 2024
Closest in time.
Predicting emergent capabilities by finetuning
C. Snell, E. Wallace, D. Klein, and S. Levine · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
G. Team · 2024
Closest in time.
Toward self-improvement of llms via imagination, searching, and criticizing, 2024
Y. Tian, B. Peng, L. Song, L. Jin, D. Yu, H. Mi, and D. Yu · 2024
Closest in time.
Trading off compute in training and inference, 2023
P. Villalobos and D. Atkinson · 2024
Closest in time.
Hypothesis search: Inductive reasoning with language models, 2024
R. Wang, E. Zelikman, G. Poesia, Y. Pu, N. Haber, and N. D. Goodman · 2024
Closest in time.
Quiet-star: Language models can teach themselves to think before speaking, 2024
E. Zelikman, G. Harik, Y. Shao, V. Jayasiri, N. Haber, and N. D. Goodman · 2024
Closest in time.