Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have significantly advanced formal theorem proving, yet the scarcity of high-quality training data constrains their capabilities in complex mathematical domains.
Combinatorial Identities
H. W. Gould · 1972
Earlier work this paper cites.
Transfer of Rule-Based Expertise through a Tutorial Dialogue
William J. Clancey · 1979
Earlier work this paper cites.
The need for biases in learning generalizations
T. M. Mitchell · 1980
Earlier work this paper cites.
New ways to make microcircuits smaller
Arthur L. Robinson · 1980
Earlier work this paper cites.
New Ways to Make Microcircuits Smaller—Duplicate Entry
Arthur L. Robinson · 1980
Earlier work this paper cites.
Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education
William J. Clancey · 1983
Earlier work this paper cites.
Strategic Explanations in Consultation—Duplicate
Diane Warner Hasling, William J. Clancey, Glenn R. Rennels, and Thomas Test · 1983
Earlier work this paper cites.
System description: E 1.8
Stephan Schulz · 1983
Earlier work this paper cites.
Classification Problem Solving
William J. Clancey · 1984
Earlier work this paper cites.
Strategic explanations for a diagnostic consultation system
Diane Warner Hasling, William J. Clancey, and Glenn Rennels · 1984
Earlier work this paper cites.
Heuristics: intelligent search strategies for computer problem solving
Judea Pearl · 1984
Earlier work this paper cites.
Blackboard Systems
Robert Engelmore and Anthony Morgan, editors · 1986
Earlier work this paper cites.
Natural deduction as higher-order resolution
Lawrence C Paulson · 1986
Earlier work this paper cites.
Poligon: A System for Parallel Problem Solving
James Rice · 1986
Earlier work this paper cites.
What is enumerative combinatorics?
Richard P Stanley · 1986
Earlier work this paper cites.
Computational Complexity of Machine Learning
M. J. Kearns · 1989
Earlier work this paper cites.
The coq proof assistant-reference manual
Projet Coq · 1996
Earlier work this paper cites.
Hol light: A tutorial introduction
John Harrison · 1996
Earlier work this paper cites.
A computer language for pure mathematics, 1997
Norman Megill · 1997
Earlier work this paper cites.
Crafting papers on machine learning
P. Langley · 2000
Earlier work this paper cites.
Combinatorial Identities
Jihuai Shi · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
E–a brainiac theorem prover
Stephan Schulz · 2002
Earlier work this paper cites.
Introductory combinatorics
Richard A Brualdi · 2004
Earlier work this paper cites.
Release of prover9
William McCune · 2005
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Z3: An efficient smt solver
Leonardo De Moura and Nikolaj Bjørner · 2008
Earlier work this paper cites.
Formal proof–the four-color theorem
Georges Gonthier et al · 2008
Earlier work this paper cites.
An acl2 tutorial
Matt Kaufmann and J Strother Moore · 2008
Earlier work this paper cites.
Proof assistants: History, ideas and future
Herman Geuvers · 2009
Earlier work this paper cites.
An Episodic History of Mathematics: Mathematical Culture through Problem Solving
SG Krantz · 2010
Earlier work this paper cites.
First-order theorem proving and vampire
Laura Kovács and Andrei Voronkov · 2013
Earlier work this paper cites.
The sketch engine: ten years on
Adam Kilgarriff, Vít Baisa, Jan Bušta, Miloš Jakubíček, Vojtěch Kovář, Jan Michelfeit, Pavel Rychlỳ, and Vít Suchomel · 2014
Earlier work this paper cites.
The lean theorem prover (system description)
Leonardo De Moura, Soonho Kong, Jeremy Avigad, Floris Van Doorn, and Jakob von Raumer · 2015
Earlier work this paper cites.
The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes
German Ros, Laura Sellart, Joanna Materzynska, David Vazquez, and Antonio M Lopez · 2016
Earlier work this paper cites.
A formal proof of the Kepler conjecture
Thomas Hales, Mark Adams, Gertrud Bauer, Tat Dat Dang, John Harrison, Hoang Le Truong, Cezary Kaliszyk, Victor Magron, Sean McLaughlin, Tat Thang Nguyen, et al · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Cezary Kaliszyk, François Chollet, and Christian Szegedy · 2017
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Draft, sketch, and prove: Guiding formal theorem provers with informal proofs
Albert Q Jiang, Sean Welleck, Jin Peng Zhou, Wenda Li, Jiacheng Liu, Mateja Jamnik, Timothée Lacroix, Yuhuai Wu, and Guillaume Lample · 2022
Later among the works it cites.
Thor: Wielding hammers to integrate language models and automated theorem provers
Albert Qiaochu Jiang, Wenda Li, Szymon Tworkowski, Konrad Czechowski, Tomasz Odrzygóźdź, Piotr Miłoś, Yuhuai Wu, and Mateja Jamnik · 2022
Later among the works it cites.
Hypertree proof search for neural theorem proving
Guillaume Lample, Timothee Lacroix, Marie-Anne Lachaux, Aurelien Rodriguez, Amaury Hayat, Thibaut Lavril, Gabriel Ebner, and Xavier Martinet · 2022
Later among the works it cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ashish Vaswani · 2017
Cited alongside, same era.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Gamepad: A learning environment for theorem proving
Daniel Huang, Prafulla Dhariwal, Dawn Song, and Ilya Sutskever · 2018
Cited alongside, same era.
Gamepad: A learning environment for theorem proving
Daniel Huang, Prafulla Dhariwal, Dawn Song, and Ilya Sutskever · 2018
Cited alongside, same era.
Pluto: The ’other’ red planet
NASA · 2018
Cited alongside, same era.
Atpboost: Learning premise selection in binary setting with atp feedback
Bartosz Piotrowski and Josef Urban · 2018
Cited alongside, same era.
Holist: An environment for machine learning of higher order logic theorem proving
Kshitij Bansal, Sarah Loos, Markus Rabe, Christian Szegedy, and Stewart Wilcox · 2019
Cited alongside, same era.
Later among the works it cites.
Synthetic proof term data augmentation for theorem proving with language models
Joseph Palermo, Johnny Ye, and Jesse Michael Han · 2022
Later among the works it cites.
Formal mathematics statement curriculum learning
Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, and Ilya Sutskever · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners, 2022
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Later among the works it cites.
Minif2f: a cross-system benchmark for formal olympiad-level mathematics, 2022
Kunhao Zheng, Jesse Michael Han, and Stanislas Polu · 2022
Later among the works it cites.
Solving math word problems via cooperative reasoning induced language models
Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Ruyi Gan, Jiaxing Zhang, and Yujiu Yang · 2022
Later among the works it cites.
Proofnet: Autoformalizing and formally proving undergraduate-level mathematics
Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W Ayers, Dragomir Radev, and Jeremy Avigad · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Later among the works it cites.
Baldur: Whole-proof generation and repair with large language models
Emily First, Markus N Rabe, Talia Ringer, and Yuriy Brun · 2023
Later among the works it cites.
Contributions to Neural Theorem Proving
Jesse Michael Han · 2023
Later among the works it cites.
Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models, 2023
Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Shuaiwen Leon Song, Samyam Rajbhandari, and Yuxiong He · 2023
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Later among the works it cites.
Proof repair infrastructure for supervised models: Building a large proof repair dataset
Tom Reichel, R Henderson, Andrew Touchet, Andrew Gardner, and Talia Ringer · 2023
Later among the works it cites.
Enhancing neural theorem proving through data augmentation and dynamic sampling method
Rahul Vishwakarma and Subhankar Mishra · 2023
Later among the works it cites.
Dt-solver: Automated theorem proving with dynamic-tree sampling guided by proof-level value function
Haiming Wang, Ye Yuan, Zhengying Liu, Jianhao Shen, Yichun Yin, Jing Xiong, Enze Xie, Han Shi, Yujun Li, Lin Li, et al · 2023
Later among the works it cites.
DT-solver: Automated theorem proving with dynamic-tree sampling guided by proof-level value function
Haiming Wang, Ye Yuan, Zhengying Liu, Jianhao Shen, Yichun Yin, Jing Xiong, Enze Xie, Han Shi, Yujun Li, Lin Li, Jian Yin, Zhenguo Li, and Xiaodan Liang · 2023
Later among the works it cites.
Lego-prover: Neural theorem proving with growing libraries
Huajian Xin, Haiming Wang, Chuanyang Zheng, Lin Li, Zhengying Liu, Qingxing Cao, Yinya Huang, Jing Xiong, Han Shi, Enze Xie, et al · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al · 2024
Later among the works it cites.
Mathstral: Advancing mathematical reasoning with 7b parameters
Mistral AI · 2024
Later among the works it cites.
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, and et al Alex Vaughan · 2024
Later among the works it cites.
Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions
Jia Li, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Huang, Kashif Rasul, Longhui Yu, Albert Q Jiang, Ziju Shen, et al · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size, 2024
Gemma Team, Shreya Pathak Morgane Riviere, and et al · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
Solving olympiad geometry without human demonstrations
Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong · 2024
Later among the works it cites.
Zijian Wu, Suozhi Huang, Zhejian Zhou, Huaiyuan Ying, Jiayu Wang, Dahua Lin, and Kai Chen · 2024
Later among the works it cites.
Deepseek-prover: Advancing theorem proving in llms through large-scale synthetic data
Huajian Xin, Daya Guo, Zhihong Shao, Zhizhou Ren, Qihao Zhu, Bo Liu, Chong Ruan, Wenda Li, and Xiaodan Liang · 2024
Later among the works it cites.
Leandojo: Theorem proving with retrieval-augmented language models
Kaiyu Yang, Aidan Swope, Alex Gu, Rahul Chalamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan J Prenger, and Animashree Anandkumar · 2024
Later among the works it cites.
Ai achieves silver medal standard solving international mathematics olympiad problems, 2024
LessWrong · 2025
Closest in time.