Fetching the paper…
Reading the bibliography…
Recent advances in automated theorem proving leverages language models to explore expanded search spaces by step-by-step proof generation.
Isabelle a Generic Theorem Prover
L. C. Paulson · 1994
Earlier work this paper cites.
The Coq Proof Assistant Reference Manual : Version 6.1
B. Barras, S. Boutin, C. Cornes, J. Courant, J.-C. Filliâtre, E. Giménez, H. Herbelin, G. Huet, C. Muñoz, C. Murthy, C. Parent, C. Paulin-Mohring, A. Saïbi, and B. Werner · 1997
Earlier work this paper cites.
Isabelle/HOL: a proof assistant for higher-order logic
T. Nipkow, M. Wenzel, and L. C. Paulson · 2002
Earlier work this paper cites.
The isabelle/isar reference manual, 2004
M. Wenzel et al · 2004
Earlier work this paper cites.
Formal verification with isabelle/hol in practice: finding a bug in the gcc scheduler
L. Gesellensetter, S. Glesner, and E. Salecker · 2008
Earlier work this paper cites.
HOL light: An overview
J. Harrison · 2009
Earlier work this paper cites.
sel4: Formal verification of an os kernel
G. Klein, K. Elphinstone, G. Heiser, J. Andronick, D. Cock, P. Derrin, D. Elkaduwe, K. Engelhardt, R. Kolanski, M. Norrish, et al · 2009
Earlier work this paper cites.
Generative language modeling for automated theorem proving
S. Polu and I. Sutskever · 2009
Earlier work this paper cites.
Three years of experience with sledgehammer, a practical link between automatic and interactive theorem provers
L. C. Paulson · 2010
Earlier work this paper cites.
The Lean Theorem Prover (System Description)
L. de Moura, S. Kong, J. Avigad, F. van Doorn, and J. von Raumer · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Imo grand challenge
R. B. e. a. Daniel Selsam, Kevin Buzzard · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy · 2020
Cited alongside, same era.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Cited alongside, same era.
The pareto principle
R. Dunford, Q. Su, and E. Tamang · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset, 2021
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Baldur: Whole-proof generation and repair with large language models
E. First, M. N. Rabe, T. Ringer, and Y. Brun · 2023
Later among the works it cites.
Fimo: A challenge formal dataset for automated theorem proving
C. Liu, J. Shen, H. Xin, Z. Liu, Y. Yuan, H. Wang, W. Ju, C. Zheng, Y. Yin, L. Li, et al · 2023
Later among the works it cites.
Magnushammer: A transformer-based approach to premise selection
M. Mikuła, S. Antoniak, S. Tworkowski, A. Q. Jiang, J. P. Zhou, C. Szegedy, Ł. Kuciński, P. Miłoś, and Y. Wu · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Q. Jiang, W. Li, J. M. Han, and Y. Wu · 2021
Cited alongside, same era.
An evaluation of the archive of formal proofs, 2021
C. MacKenzie, J. Fleuriot, and J. Vaughan · 2021
Cited alongside, same era.
miniF2F: a cross-system benchmark for formal Olympiad-level mathematics
K. Zheng, J. M. Han, and S. Polu · 2021
Cited alongside, same era.
Proof artifact co-training for theorem proving with language models
J. M. Han, J. Rute, Y. Wu, E. W. Ayers, and S. Polu · 2022
Cited alongside, same era.
HyperTree Proof Search for Neural Theorem Proving
G. Lample, M.-A. Lachaux, T. Lavril, X. Martinet, A. Hayat, G. Ebner, A. Rodriguez, and T. Lacroix · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra · 2022
Cited alongside, same era.
Formal Mathematics Statement Curriculum Learning
S. Polu, J. M. Han, K. Zheng, M. Baksys, I. Babuschkin, and I. Sutskever · 2022
Cited alongside, same era.
Autoformalization with large language models
Y. Wu, A. Q. Jiang, W. Li, M. Rabe, C. Staats, M. Jamnik, and C. Szegedy · 2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2023
Later among the works it cites.
Trigo: Benchmarking formal mathematical proof reduction for generative language models, 2023
J. Xiong, J. Shen, Y. Yuan, H. Wang, Y. Yin, Z. Liu, L. Li, Z. Guo, Q. Cao, Y. Huang, C. Zheng, X. Liang, M. Zhang, and Q. Liu · 2023
Later among the works it cites.
Leandojo: Theorem proving with retrieval-augmented language models
K. Yang, A. M. Swope, A. Gu, R. Chalamala, P. Song, S. Yu, S. Godil, R. Prenger, and A. Anandkumar · 2023
Later among the works it cites.
Lyra: Orchestrating dual correction in automated theorem proving, 2023
C. Zheng, H. Wang, E. Xie, Z. Liu, J. Sun, H. Xin, J. Shen, Z. Li, and Y. Li · 2023
Later among the works it cites.
Subgoal search for complex reasoning tasks, 2024
K. Czechowski, T. Odrzygóźdź, M. Zbysiński, M. Zawalski, K. Olejnik, Y. Wu, Łukasz Kuciński, and P. Miłoś · 2024
Closest in time.
Mustard: Mastering uniform synthesis of theorem and proof data, 2024
Y. Huang, X. Lin, Z. Liu, Q. Cao, H. Xin, H. Wang, Z. Li, L. Song, and X. Liang · 2024
Closest in time.
An in-context learning agent for formal theorem-proving, 2024
A. Thakur, G. Tsoukalas, Y. Wen, J. Xin, and S. Chaudhuri · 2024
Closest in time.
Selene: Pioneering automated proof in software verification, 2024
L. Zhang, S. Lu, and N. Duan · 2024
Closest in time.