Fetching the paper…
Reading the bibliography…
Recent advancements in Large Language Models (LLMs) have created new opportunities to enhance performance on complex reasoning tasks by leveraging test-time computation.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Making large language models better reasoners with step-aware verifier
Y. Li, Z. Lin, S. Zhang, Q. Fu, B. Chen, J.-G. Lou, and W. Chen · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Earlier work this paper cites.
Langfun, Sept. 2023
D. Peng · 2023
Earlier work this paper cites.
Self-evaluation improves selective generation in large language models
J. Ren, Y. Zhao, T. Vu, P. J. Liu, and B. Lakshminarayanan · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Earlier work this paper cites.
Large language models are better reasoners with self-verification
Y. Weng, M. Zhu, F. Xia, B. Li, S. He, S. Liu, B. Sun, K. Liu, and J. Zhao · 2023
Earlier work this paper cites.
Claude 3.5 Sonnet
anthropic · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini · 2024
Cited alongside, same era.
Are more llm calls all you need? towards the scaling properties of compound ai systems
L. Chen, J. Q. Davis, B. Hanin, P. Bailis, I. Stoica, M. Zaharia, and J. Zou · 2024
Cited alongside, same era.
Ticking all the boxes: Generated checklists improve llm evaluation and generation
J. Cook, T. Rocktäschel, J. Foerster, D. Aumiller, and A. Wang · 2024
Cited alongside, same era.
T. P. Ferraz, K. Mehta, Y.-H. Lin, H.-S. Chang, S. Oraby, S. Liu, V. Subramanian, T. Chung, M. Bansal, and N. Peng · 2024
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al · 2024
Later among the works it cites.
From decoding to meta-generation: Inference-time algorithms for large language models
S. Welleck, A. Bertsch, M. Finlayson, H. Schoelkopf, A. Xie, G. Neubig, I. Kulikov, and Z. Harchaoui · 2024
Later among the works it cites.
Livebench: A challenging, contamination-free llm benchmark
C. White, S. Dooley, M. Roberts, A. Pal, B. Feuer, S. Jain, R. Shwartz-Ziv, N. Jain, K. Saifullah, S. Naidu, C. Hegde, Y. LeCun, T. Goldstein, W. Neiswanger, and M. Goldblum · 2024
Later among the works it cites.
Y. Wu, Z. Sun, S. Li, S. Welleck, and Y. Yang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica · 2024
Cited alongside, same era.
Improving llm reasoning through scaling inference computation with collaborative verification
Z. Liang, Y. Liu, T. Niu, X. Zhang, Y. Zhou, and S. Yavuz · 2024
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al · 2024
Cited alongside, same era.
Recursive introspection: Teaching language model agents how to self-improve
Y. Qu, T. Zhang, N. Garg, and A. Kumar · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Cited alongside, same era.
Mind the gap: Examining the self-improvement capabilities of large language models
Y. Song, H. Zhang, C. Eisenach, S. Kakade, D. Foster, and U. Ghai · 2024
Cited alongside, same era.
Critic: Large language models can self-correct with tool-interactive critiquing
Z. Gou, Z. Shao, Y. Gong, Y. Yang, N. Duan, W. Chen, et al
Cited in the paper.
Scaling llm inference with optimized sample compute allocation
K. Zhang, S. Zhou, D. Wang, W. Y. Wang, and L. Li · 2024
Later among the works it cites.
Natural plan: Benchmarking llms on natural language planning
H. S. Zheng, S. Mishra, H. Zhang, X. Chen, M. Chen, A. Nova, L. Hou, H.-T. Cheng, Q. V. Le, E. H. Chi, et al · 2024
Later among the works it cites.
https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
American invitational mathematics examination · 2025
Closest in time.
https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/
Gemini 2.5: Our most intelligent ai model · 2025
Closest in time.
Inference-time scaling for complex tasks: Where we stand and what lies ahead
V. Balachandran, J. Chen, L. Chen, S. Garg, N. Joshi, Y. Lara, J. Langford, B. Nushi, V. Vineet, Y. Wu, et al · 2025
Closest in time.
K.-H. Lee, I. Fischer, Y.-H. Wu, D. Marwood, S. Baluja, D. Schuurmans, and X. Chen · 2025
Closest in time.
Sample, scrutinize and scale: Effective inference-time search by scaling verification
E. Zhao, P. Awasthi, and S. Gollapudi · 2025
Closest in time.