Fetching the paper…
Reading the bibliography…
Inference-time scaling can enhance the reasoning capabilities of large language models (LLMs) on complex problems that benefit from step-by-step problem solving.
Computers and intractability: a guide to the theory of np-completeness (michael r. garey and david s. johnson)
Juris Hartmanis · 1982
Earlier work this paper cites.
Sample, scrutinize and scale: Effective inference-time search by scaling verification, 2025
Eric Zhao, Pranjal Awasthi, and Sreenivas Gollapudi · 1983
Earlier work this paper cites.
Where the really hard problems are
Peter C Cheeseman, Bob Kanefsky, William M Taylor, et al · 1991
Earlier work this paper cites.
Hard and easy distributions of sat problems
David Mitchell, Bart Selman, Hector Levesque, et al · 1992
Earlier work this paper cites.
Critical behavior in the satisfiability of random boolean expressions
Scott Kirkpatrick and Bart Selman · 1994
Earlier work this paper cites.
Computational complexity
Christos H Papadimitriou · 2003
Earlier work this paper cites.
Random satisfiability
Dimitris Achlioptas · 2009
Earlier work this paper cites.
Reducibility among combinatorial problems
Richard M Karp · 2009
Earlier work this paper cites.
Overreliance on ai literature review
Samir Passi and Mihaela Vorvoreanu · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Earlier work this paper cites.
Reinforced self-training (rest) for language modeling
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Earlier work this paper cites.
Making language models better reasoners with step-aware verifier
Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen · 2023
Earlier work this paper cites.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Cited alongside, same era.
Tree of thoughts: Deliberate problem solving with large language models, 2023
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Cited alongside, same era.
Claude 3.5 sonnet
Anthropic · 2024
Cited alongside, same era.
Eureka: Evaluating and understanding large foundation models
Vidhisha Balachandran, Jingya Chen, Neel Joshi, Besmira Nushi, Hamid Palangi, Eduardo Salinas, Vibhav Vineet, James Woffinden-Luey, and Safoora Yousefi · 2024
Cited alongside, same era.
Diversity of thought improves reasoning abilities of llms, 2024
Ranjita Naik, Varun Chandrasekaran, Mert Yuksekgonul, Hamid Palangi, and Besmira Nushi · 2024
Later among the works it cites.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman · 2024
Later among the works it cites.
Scaling test-time compute optimally can be more effective than scaling llm parameters
Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
From decoding to meta-generation: Inference-time algorithms for large language models
Sean Welleck, Amanda Bertsch, Matthew Finlayson, Hailey Schoelkopf, Alex Xie, Graham Neubig, Ilia Kulikov, and Zaid Harchaoui · 2024
Later among the works it cites.
Aime 83-24
AIME · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large language monkeys: Scaling inference compute with repeated sampling
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini · 2024
Cited alongside, same era.
Benchagents: Automated benchmark creation with agent interaction
Natasha Butt, Varun Chandrasekaran, Neel Joshi, Besmira Nushi, and Vidhisha Balachandran · 2024
Cited alongside, same era.
Are more llm calls all you need? towards the scaling properties of compound ai systems
Lingjiao Chen, Jared Davis, Boris Hanin, Peter Bailis, Ion Stoica, Matei Zaharia, and James Zou · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Cited alongside, same era.
Nphardeval: Dynamic benchmark on reasoning ability of large language models via complexity classes
Lizhou Fan, Wenyue Hua, Lingyao Li, Haoyang Ling, and Yongfeng Zhang · 2024
Cited alongside, same era.
Can large language models reason? a characterization via 3-sat
Rishi Hazra, Gabriele Venturato, Pedro Zuidberg Dos Martires, and Luc De Raedt · 2024
Cited alongside, same era.
V-star: Training verifiers for self-taught reasoners
Arian Hosseini, Xingdi Yuan, Nikolay Malkin, Aaron Courville, Alessandro Sordoni, and Rishabh Agarwal · 2024
Cited alongside, same era.
AIME · 2025
Closest in time.
Claude 3.7 sonnet
Anthropic · 2025
Closest in time.
A taxonomy of linguistic expressions that contribute to anthropomorphism of language technologies
Alicia DeVrio, Myra Cheng, Lisa Egede, Alexandra Olteanu, and Su Lin Blodgett · 2025
Closest in time.
Omni-math: A universal olympiad level mathematic benchmark for large language models
Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, et al · 2025
Closest in time.
Gemini 2.0 pro experimental
Google · 2025
Closest in time.
Gemini 2.0 flash thinking
Google · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.
Openai o3-mini system card
OpenAI · 2025
Closest in time.
Inference scaling laws: An empirical analysis of compute-optimal inference for llm problem-solving
Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang · 2025
Closest in time.