Fetching the paper…
Reading the bibliography…
Solving mathematical problems using computer-verifiable languages like Lean has significantly impacted the mathematical and computer science communities.
The logic theory machine–a complex information processing system
Newell, A. and Simon, H · 1956
Earlier work this paper cites.
Isabelle: A generic theorem prover
Paulson, L. C · 1994
Earlier work this paper cites.
Hol light: An overview
Harrison, J · 2009
Earlier work this paper cites.
The lean theorem prover (system description)
De Moura, L., Kong, S., Avigad, J., Van Doorn, F., and von Raumer, J · 2015
Earlier work this paper cites.
Cem-rl: Combining evolutionary and gradient-based methods for policy search
Pourchot, A. and Sigaud, O · 2018
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Polu, S. and Sutskever, I · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Lisa: Language models of isabelle proofs
Jiang, A. Q., Li, W., Han, J. M., and Wu, Y · 2021
Earlier work this paper cites.
The lean 4 theorem prover and programming language
Moura, L. d. and Ullrich, S · 2021
Earlier work this paper cites.
Minif2f: a cross-system benchmark for formal olympiad-level mathematics
Zheng, K., Han, J. M., and Polu, S · 2021
Earlier work this paper cites.
Draft, sketch, and prove: Guiding formal theorem provers with informal proofs
Jiang, A. Q., Welleck, S., Zhou, J. P., Li, W., Liu, J., Jamnik, M., Lacroix, T., Wu, Y., and Lample, G · 2022
Earlier work this paper cites.
Formal mathematics statement curriculum learning
Polu, S., Han, J. M., Zheng, K., Baksys, M., Babuschkin, I., and Sutskever, I · 2022
Earlier work this paper cites.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al · 2022
Cited alongside, same era.
Scienceworld: Is your agent smarter than a 5th grader?
Wang, R., Jansen, P., Côté, M.-A., and Ammanabrolu, P · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Leanagent: Lifelong learning for formal theorem proving
Kumarappan, A., Tiwari, M., Song, P., George, R. J., Xiao, C., and Anandkumar, A · 2024
Later among the works it cites.
Lean-star: Learning to interleave thinking and proving
Lin, H., Sun, Z., Yang, Y., and Welleck, S · 2024
Later among the works it cites.
Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al · 2024
Later among the works it cites.
Agentboard: An analytical evaluation board of multi-turn llm agents
Ma, C., Zhang, J., Zhu, Z., Yang, C., Yang, Y., Jin, Y., Lan, Z., Kong, L., and He, J · 2024
Later among the works it cites.
Open-o1, 2024
Open-Source-O1 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S., Jiang, A. Q., Deng, J., Biderman, S., and Welleck, S · 2023
Cited alongside, same era.
Baldur: Whole-proof generation and repair with large language models
First, E., Rabe, M. N., Ringer, T., and Brun, Y · 2023
Cited alongside, same era.
Wang, R., Zhou, W., and Sachan, M · 2023
Cited alongside, same era.
Openagents: An open platform for language agents in the wild
Xie, T., Zhou, F., Cheng, Z., Shi, P., Weng, L., Liu, Y., Hua, T. J., Zhao, J., Liu, Q., Liu, C., et al · 2023
Cited alongside, same era.
Lemur: Harmonizing natural language and code for language agents
Xu, Y., Su, H., Xing, C., Mi, B., Liu, Q., Shi, W., Hui, B., Zhou, F., Liu, Y., Xie, T., et al · 2023
Cited alongside, same era.
Auto-gpt for online decision making: Benchmarks and additional opinions
Yang, H., Yue, S., and He, Y · 2023
Cited alongside, same era.
Lyra: Orchestrating dual correction in automated theorem proving
Zheng, C., Wang, H., Xie, E., Liu, Z., Sun, J., Xin, H., Shen, J., Li, Z., and Li, Y · 2023
Cited alongside, same era.
Data for mathematical copilots: Better ways of presenting proofs for machine learning
Frieder, S., Bayer, J., Collins, K. M., Berner, J., Loader, J., Juhász, A., Ruehle, F., Welleck, S., Poesia, G., Griffiths, R.-R., et al · 2024
Cited alongside, same era.
Later among the works it cites.
Learning to reason with llms
OpenAI · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Later among the works it cites.
Theoremllama: Transforming general-purpose llms into lean4 experts
Wang, R., Zhang, J., Jia, Y., Pan, R., Diao, S., Pi, R., and Zhang, T · 2024
Later among the works it cites.
Lean workbook: A large-scale lean problem set formalized from natural language math problems
Ying, H., Wu, Z., Geng, Y., Wang, J., Lin, D., and Chen, K · 2024
Later among the works it cites.
Beyond limited data: Self-play llm theorem provers with iterative conjecturing and proving
Dong, K. and Ma, T · 2025
Closest in time.
Goedel-prover: A frontier model for open-source automated theorem proving, 2025
Lin, Y., Tang, S., Lyu, B., Wu, J., Lin, H., Yang, K., Li, J., Xia, M., Chen, D., Arora, S., and Jin, C · 2025
Closest in time.