Fetching the paper…
Reading the bibliography…
Neurosymbolic approaches integrating large language models with formal reasoning have recently achieved human-level performance on mathematics competition problems in algebra, geometry and number theory.
A deductive database approach to automated geometry theorem proving and discovering
S.-C. Chou, X.-S. Gao, and J.-Z. Zhang · 2000
Earlier work this paper cites.
E–a brainiac theorem prover
S. Schulz · 2002
Earlier work this paper cites.
Introductory combinatorics
R. A. Brualdi · 2004
Earlier work this paper cites.
The isabelle framework
M. Wenzel, L. C. Paulson, and T. Nipkow · 2008
Earlier work this paper cites.
Automated theorem proving
W. Bibel · 2013
Earlier work this paper cites.
The Lean theorem prover (system description)
L. de Moura, S. Kong, J. Avigad, F. Van Doorn, and J. von Raumer · 2015
Earlier work this paper cites.
Minif2f: a cross-system benchmark for formal olympiad-level mathematics
K. Zheng, J. M. Han, and S. Polu · 2017
Earlier work this paper cites.
The lean mathematical library
T. mathlib Community · 2020
Earlier work this paper cites.
Generative language modeling for automated theorem proving
S. Polu and I. Sutskever · 2020
Earlier work this paper cites.
The lean 4 theorem prover and programming language
L. d. Moura and S. Ullrich · 2021
Earlier work this paper cites.
Draft, sketch, and prove: Guiding formal theorem provers with informal proofs
A. Q. Jiang, S. Welleck, J. P. Zhou, W. Li, J. Liu, M. Jamnik, T. Lacroix, Y. Wu, and G. Lample · 2022
Cited alongside, same era.
Proofnet: Autoformalizing and formally proving undergraduate-level mathematics
Z. Azerbayev, B. Piotrowski, H. Schoelkopf, E. W. Ayers, D. Radev, and J. Avigad · 2023
Cited alongside, same era.
Fimo: A challenge formal dataset for automated theorem proving, 2023
C. Liu, J. Shen, H. Xin, Z. Liu, Y. Yuan, H. Wang, W. Ju, C. Zheng, Y. Yin, L. Li, M. Zhang, and Q. Liu · 2023
Cited alongside, same era.
The Coq Proof Assistant, Sept. 2023
The Coq Development Team · 2023
Cited alongside, same era.
IMO Grand Challenge — imo-grand-challenge.github.io
I. G. Challenge · 2024
Cited alongside, same era.
Gemini 2.5: Our most intelligent ai model, 2025
Google · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Goedel-prover: A frontier model for open-source automated theorem proving
Y. Lin, S. Tang, B. Lyu, J. Wu, H. Lin, K. Yang, J. Li, M. Xia, D. Chen, S. Arora, et al · 2025
Closest in time.
Openai o3-mini, 2025
OpenAI · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning, March 2025
T. Qwen · 2025
Closest in time.
Kimina-prover preview: Towards large formal reasoning models with reinforcement learning, 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Cited alongside, same era.
AIMO Prize — aimoprize.com
Prize · 2024
Cited alongside, same era.
Putnambench: Evaluating neural theorem-provers on the putnam mathematical competition
G. Tsoukalas, J. Lee, J. Jennings, J. Xin, M. Ding, M. Jennings, A. Thakur, and S. Chaudhuri · 2024
Cited alongside, same era.
Claude 3.7 sonnet, 2025
Anthropic · 2025
Cited alongside, same era.
Stp: Self-play llm theorem provers with iterative conjecturing and proving
K. Dong and T. Ma · 2025
Cited alongside, same era.
Lego-prover: Neural theorem proving with growing libraries, 2023a
H. Wang, H. Xin, C. Zheng, L. Li, Z. Liu, Q. Cao, Y. Huang, J. Xiong, H. Shi, E. Xie, J. Yin, Z. Li, H. Liao, and X. Liang
Cited in the paper.
Dt-solver: Automated theorem proving with dynamic-tree sampling guided by proof-level value function
H. Wang, Y. Yuan, Z. Liu, J. Shen, Y. Yin, J. Xiong, E. Xie, H. Shi, Y. Li, L. Li, et al
Cited in the paper.
H. Wang, M. Unsal, X. Lin, M. Baksys, J. Liu, M. D. Santos, F. Sung, M. Vinyes, Z. Ying, Z. Zhu, J. Lu, H. de Saxcé, B. Bailey, C. Song, C. Xiao, D. Zhang, E. Zhang, F. Pu, H. Zhu, J. Liu, J. Bayer, J. Michel, L. Yu, L. Dreyfus-Schmidt, L. Tunstall, L. Pagani, M. Machado, P. Bourigault, R. Wang, S. Polu, T. Barroyer, W.-D. Li, Y. Niu, Y. Fleureau, Y. Hu, Z. Yu, Z. Wang, Z. Yang, Z. Liu, and J. Li · 2025
Closest in time.
Bfs-prover: Scalable best-first tree search for llm-based automatic theorem proving
R. Xin, C. Xi, J. Yang, F. Chen, H. Wu, X. Xiao, Y. Sun, S. Zheng, and K. Shen · 2025
Closest in time.
A combinatorial identities benchmark for theorem proving via automated theorem generation
B. Xiong, H. Lv, H. Shan, J. Wang, Z. Yang, and L. Zhi · 2025
Closest in time.
Leanabell-prover: Posttraining scaling in formal reasoning
J. Zhang, Q. Wang, X. Ji, Y. Liu, Y. Yue, F. Zhang, D. Zhang, G. Zhou, and K. Gai · 2025
Closest in time.