Fetching the paper…
Reading the bibliography…
Formal mathematical reasoning remains a critical challenge for artificial intelligence, hindered by limitations of existing benchmarks in scope and scale.
Problems in Mathematical Analysis. Edited by B. Demidovich. Translated From the Russian by G. Yankovsky
B.P. Demidovich · 1964
Earlier work this paper cites.
Constructive mathematics
Brian Smith · 1995
Earlier work this paper cites.
The Coq proof assistant reference manual: Version 6.1
Bruno Barras, Samuel Boutin, Cristina Cornes, Judicaël Courant, Jean-Christophe Filliatre, Eduardo Gimenez, Hugo Herbelin, Gerard Huet, Cesar Munoz, Chetan Murthy, et al · 1997
Earlier work this paper cites.
Isabelle/HOL: a proof assistant for higher-order logic
Tobias Nipkow, Markus Wenzel, and Lawrence C Paulson · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Efficient selectivity and backup operators in Monte-Carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Bandit based Monte-Carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Dafny: An automatic program verifier for functional correctness
K Rustan M Leino · 2010
Earlier work this paper cites.
Problems and solutions in real analysis
Masayoshi Hata · 2016
Earlier work this paper cites.
The Lean mathematical library
Mathlib Community · 2020
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
The lean 4 theorem prover and programming language
Leonardo de Moura and Sebastian Ullrich · 2021
Earlier work this paper cites.
Minif2f: a cross-system benchmark for formal olympiad-level mathematics
Kunhao Zheng, Jesse Michael Han, and Stanislas Polu · 2021
Earlier work this paper cites.
Liquid tensor experiment
Peter Scholze · 2022
Earlier work this paper cites.
Autoformalization with large language models
Yuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus Rabe, Charles Staats, Mateja Jamnik, and Christian Szegedy · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Proofnet: Autoformalizing and formally proving undergraduate-level mathematics
Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W Ayers, Dragomir Radev, and Jeremy Avigad · 2023
Cited alongside, same era.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Cited alongside, same era.
A read-eval-print-loop for Lean 4
Leanprover Community · 2023
Cited alongside, same era.
Aesop: White-box best-first proof search for lean
Jannis Limperg and Asta Halkjær From · 2023
Cited alongside, same era.
Fimo: A challenge formal dataset for automated theorem proving
Chengwu Liu, Jianhao Shen, Huajian Xin, Zhengying Liu, Ye Yuan, Haiming Wang, Wei Ju, Chuanyang Zheng, Yichun Yin, Lin Li, et al · 2023
Cited alongside, same era.
Zijian Wu, Suozhi Huang, Zhejian Zhou, Huaiyuan Ying, Jiayu Wang, Dahua Lin, and Kai Chen · 2024
Later among the works it cites.
Verbalized machine learning: Revisiting machine learning with language models
Tim Z Xiao, Robert Bamler, Bernhard Schölkopf, and Weiyang Liu · 2024
Later among the works it cites.
Huajian Xin, ZZ Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, et al · 2024
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OpenAI · 2023
Cited alongside, same era.
The polynomial freiman-ruzsa conjecture
Terence Tao · 2023
Cited alongside, same era.
LeanDojo: Theorem proving with retrieval-augmented language models
Kaiyu Yang, Aidan Swope, Alex Gu, Rahul Chalamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan Prenger, and Anima Anandkumar · 2023
Cited alongside, same era.
Ai achieves silver-medal standard solving international 178 mathematical olympiad problems
Team AlphaProof and Team AlphaGeometry · 2024
Cited alongside, same era.
U-math: A university-level benchmark for evaluating mathematical skills in llms
Konstantin Chernyshev, Vitaliy Polshkov, Ekaterina Artemova, Alex Myasnikov, Vlad Stepanov, Alexei Miasnikov, and Sergei Tilga · 2024
Cited alongside, same era.
Hardmath: A benchmark dataset for challenging problems in applied mathematics
Jingxuan Fan, Sarah Martinson, Erik Y Wang, Kaylie Hausknecht, Jonah Brenner, Danxian Liu, Nianli Peng, Corey Wang, and Michael P Brenner · 2024
Cited alongside, same era.
Omni-math: A universal olympiad level mathematic benchmark for large language models
Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, et al · 2024
Cited alongside, same era.
Kaiyu Yang, Gabriel Poesia, Jingxuan He, Wenda Li, Kristin Lauter, Swarat Chaudhuri, and Dawn Song · 2024
Later among the works it cites.
Bluemo: A comprehensive collection of challenging mathematical olympiad problems from the little blue book series., 2024
Yifan Zhang, Yifan Luo, and Yizhou Chen · 2024
Later among the works it cites.
Beyond limited data: Self-play llm theorem provers with iterative conjecturing and proving
Kefan Dong and Tengyu Ma · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Goedel-prover: A frontier model for open-source automated theorem proving
Yong Lin, Shange Tang, Bohan Lyu, Jiayun Wu, Hongzhou Lin, Kaiyu Yang, Jia Li, Mengzhou Xia, Danqi Chen, Sanjeev Arora, et al · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.
ZZ Ren, Zhihong Shao, Junxiao Song, Huajian Xin, Haocheng Wang, Wanjia Zhao, Liyue Zhang, Zhe Fu, Qihao Zhu, Dejian Yang, et al · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
De morgan’s laws — Wikipedia, the free encyclopedia, [Year of specific version, e.g., 2025]
Wikipedia contributors · 2025
Closest in time.
Law of excluded middle — Wikipedia, the free encyclopedia, [Year of specific version, e.g., 2025]
Wikipedia contributors · 2025
Closest in time.
Kimina-prover preview: Towards large formal reasoning models with reinforcement learning
Haiming Wang, Mert Unsal, Xiaohan Lin, Mantas Baksys, Junqi Liu, Marco Dos Santos, Flood Sung, Marina Vinyes, Zhenzhe Ying, Zekai Zhu, et al · 2025
Closest in time.
Bfs-prover: Scalable best-first tree search for llm-based automatic theorem proving
Ran Xin, Chenguang Xi, Jie Yang, Feng Chen, Hang Wu, Xia Xiao, Yifan Sun, Shen Zheng, and Kai Shen · 2025
Closest in time.
Generating symbolic world models via test-time scaling of large language models
Zhouliang Yu, Yuhuan Yuan, Tim Z Xiao, Fuxiang Frank Xia, Jie Fu, Ge Zhang, Ge Lin, and Weiyang Liu · 2025
Closest in time.