Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have shown great potential in reasoning tasks through test-time scaling methods like self-consistency with majority voting.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Let’s sample step by step: Adaptive-consistency for efficient reasoning and coding with llms
Pranjal Aggarwal, Aman Madaan, Yiming Yang, et al · 2023
Earlier work this paper cites.
Rewarding chatbots for real-world engagement with millions of users
Robert Irvine, Douglas Boubert, Vyas Raina, Adian Liusie, Ziyi Zhu, Vineet Mudupalli, Aliaksei Korshuk, Zongyi Liu, Fritz Cremer, Valentin Assassi, et al · 2023
Earlier work this paper cites.
Self-Evaluation Improves Selective Generation in Large Language Models
Jie Ren, Yao Zhao, Tu Vu, Peter J. Liu, and Balaji Lakshminarayanan · 2023
Earlier work this paper cites.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed H Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Earlier work this paper cites.
Dynamic Voting for Efficient Reasoning in Large Language Models
Mingfeng Xue, Dayiheng Liu, Wenqiang Lei, Xingzhang Ren, Baosong Yang, Jun Xie, Yidan Zhang, Dezhong Peng, and Jiancheng Lv · 2023
Earlier work this paper cites.
Large language monkeys: Scaling inference compute with repeated sampling
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini · 2024
Earlier work this paper cites.
Learning to Route LLMs with Confidence Tokens
Yu-Neng Chuang, Prathusha K. Sarma, Parikshit Gopalan, John Boccio, Sara Bolouki, Xia Hu, and Helen Zhou · 2024
Earlier work this paper cites.
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification
Ekaterina Fadeeva, Aleksandr Rubashevskii, Artem Shelmanov, Sergey Petrakov, Haonan Li, Hamdy Mubarak, Evgenii Tsymbalov, Gleb Kuzmin, Alexander Panchenko, Timothy Baldwin, Preslav Nakov, and Maxim Panov · 2024
Earlier work this paper cites.
Detecting hallucinations in large language models using semantic entropy
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal · 2024
Earlier work this paper cites.
Efficiently scaling llm reasoning with certaindex
Yichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu, Zhongdongming Dai, Yonghao Zhuang, Yian Ma, Aurick Qiao, Tajana Rosing, Ion Stoica, et al · 2024
Earlier work this paper cites.
A survey of confidence estimation and calibration in large language models
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych · 2024
Earlier work this paper cites.
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
Zhijian Han, Zhi Li, Yutong Wang, Chenglu Guo, Ruifei Song, Jiacheng He, Jun Pan, Jianing Wu, Shuai Li, Liang Gao, and Weihang Chen · 2024
Earlier work this paper cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Earlier work this paper cites.
Lightweight reranking for language model generations
Siddhartha Jain, Xiaofei Ma, Anoop Deoras, and Bing Xiang · 2024
Earlier work this paper cites.
Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning
Yiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Xinglin Wang, Bin Sun, Heda Wang, and Kan Li · 2024
Earlier work this paper cites.
Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems
Yingqian Min, Zhipeng Chen, Jinhao Jiang, Jie Chen, Jia Deng, Yiwen Hu, Yiru Tang, Jiapeng Wang, Xiaoxue Cheng, Huatong Song, et al · 2024
Cited alongside, same era.
Hint Marginalization for Improved Reasoning in Large Language Models
Soumyasundar Pal, Didier Chételat, Yingxue Zhang, and Mark Coates · 2024
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Don’t overthink it. preferring shorter thinking chains for improved llm reasoning
Michael Hassid, Gabriel Synnaeve, Yossi Adi, and Roy Schwartz · 2025
Closest in time.
Hmmt 2025
HMMT · 2025
Closest in time.
Thinkprune: Pruning long chain-of-thought of llms via reinforcement learning
Bairu Hou, Yang Zhang, Jiabao Ji, Yujian Liu, Kaizhi Qian, Jacob Andreas, and Shiyu Chang · 2025
Closest in time.
Scalable best-of-n selection for large language models via self-certainty, 2025
Zhewei Kang, Xuandong Zhao, and Dawn Song · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vernon Y.H. Toh, Deepanway Ghosal, and Soujanya Poria · 2024
Cited alongside, same era.
Soft Self-Consistency Improves Language Model Agents
Han Wang, Archiki Prasad, Elias Stengel-Eskin, and Mohit Bansal · 2024
Cited alongside, same era.
From decoding to meta-generation: Inference-time algorithms for large language models
Sean Welleck, Amanda Bertsch, Matthew Finlayson, Hailey Schoelkopf, Alex Xie, Graham Neubig, Ilia Kulikov, and Zaid Harchaoui · 2024
Cited alongside, same era.
Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang · 2024
Cited alongside, same era.
https://www.brumo.org/ , 2025
Brumo. brown university math olympiad 2025 · 2025
Cited alongside, same era.
2024 aime i
Art of Problem Solving · 2025
Cited alongside, same era.
2024 aime ii
Art of Problem Solving · 2025
Cited alongside, same era.
2025 aime i
Art of Problem Solving · 2025
Cited alongside, same era.
Closest in time.
O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning
Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, and Dacheng Tao · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.
Introducing gpt-5, 2025
OpenAI · 2025
Closest in time.
gpt-oss-120b & gpt-oss-20b model card
OpenAI · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
Reasoning aware self-consistency: Leveraging reasoning paths for efficient LLM sampling
Guangya Wan, Yuqi Wu, Jie Chen, and Sheng Li · 2025
Closest in time.
Ranked voting based self-consistency of large language models
Weiqin Wang, Yile Wang, and Hui Huang · 2025
Closest in time.
Grok 4, 2025
xAI · 2025
Closest in time.
Limo: Less is more for reasoning
Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, and Pengfei Liu · 2025
Closest in time.
Learning to Reason without External Rewards
Xuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine, and Dawn Song · 2025
Closest in time.