Fetching the paper…
Reading the bibliography…
Neural theorem proving combines large language models (LLMs) with proof assistants such as Lean, where the correctness of formal proofs can be rigorously verified, leaving no room for hallucination.
Learning to reason in large theories without imitation
Kshitij Bansal, Christian Szegedy, Markus N Rabe, Sarah M Loos, and Viktor Toman · 1905
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering, 2020
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen tau Yih · 2004
Earlier work this paper cites.
Proving equalities in a commutative ring done right in Coq
Benjamin Grégoire and Assia Mahboubi · 2005
Earlier work this paper cites.
Fast reflexive arithmetic tactics the linear case and beyond
Frédéric Besson · 2007
Earlier work this paper cites.
Sledgehammer: judgement day
Sascha Böhme and Tobias Nipkow · 2010
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Earlier work this paper cites.
The Lean theorem prover (system description)
Leonardo de Moura, Soonho Kong, Jeremy Avigad, Floris Van Doorn, and Jakob von Raumer · 2015
Earlier work this paper cites.
Hammering towards QED
Jasmin Christian Blanchette, Cezary Kaliszyk, Lawrence C Paulson, and Josef Urban · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
DeepMath—deep sequence models for premise selection
Geoffrey Irving, Christian Szegedy, Alexander A Alemi, Niklas Eén, François Chollet, and Josef Urban · 2016
Earlier work this paper cites.
Holophrasm: a neural automated theorem prover for higher-order logic, 2016
Daniel Whalen · 2016
Earlier work this paper cites.
Deepmath - deep sequence models for premise selection, 2017
Alex A. Alemi, Francois Chollet, Niklas Een, Geoffrey Irving, Christian Szegedy, and Josef Urban · 2017
Earlier work this paper cites.
SMTCoq: A plug-in for integrating SMT solvers into Coq
Burak Ekici, Alain Mebsout, Cesare Tinelli, Chantal Keller, Guy Katz, Andrew Reynolds, and Clark Barrett · 2017
Earlier work this paper cites.
HolStep: A machine learning dataset for higher-order logic theorem proving
Cezary Kaliszyk, François Chollet, and Christian Szegedy · 2017
Earlier work this paper cites.
Libnpy: a simple c++ library for reading and writing of numpy’s .npy files., 2017
Leon Merten Lohse · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Premise selection for theorem proving by deep graph embedding
Mingzhe Wang, Yihe Tang, Jian Wang, and Jia Deng · 2017
Earlier work this paper cites.
Hammer for Coq: Automation for dependent type theory
Łukasz Czajka and Cezary Kaliszyk · 2018
Earlier work this paper cites.
GamePad: A learning environment for theorem proving
Daniel Huang, Prafulla Dhariwal, Dawn Song, and Ilya Sutskever · 2019
Earlier work this paper cites.
Meta-f*: Proof automation with smt, tactics, and metaprograms, 2019
Guido Martínez, Danel Ahman, Victor Dumitrescu, Nick Giannarakis, Chris Hawblitzel, Catalin Hritcu, Monal Narasimhamurthy, Zoe Paraskevopoulou, Clément Pit-Claudel, Jonathan Protzenko, Tahina Ramananandro, Aseem Rastogi, and Nikhil Swamy · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
Learning to prove theorems via interacting with proof assistants
Kaiyu Yang and Jia Deng · 2019
Earlier work this paper cites.
CTranslate2: a c++ and python library for efficient inference with transformer models
The OpenNMT Authors · 2020
Earlier work this paper cites.
Mathematics in Lean, 2020
Jeremy Avigad and Patrick Massot · 2020
Earlier work this paper cites.
The Tactician: A seamless, interactive tactic learner and prover for Coq
Lasse Blaauwbroek, Josef Urban, and Herman Geuvers · 2020
Earlier work this paper cites.
TacTok: semantics-aware proof synthesis
Emily First, Yuriy Brun, and Arjun Guha · 2020
Earlier work this paper cites.
The Lean mathematical library
The mathlib Community · 2020
Earlier work this paper cites.
Graph representations for higher-order logic and theorem proving
Aditya Paliwal, Sarah Loos, Markus Rabe, Kshitij Bansal, and Christian Szegedy · 2020
Cited alongside, same era.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever · 2020
Cited alongside, same era.
Learning to prove theorems by learning to generate theorems
Mingzhe Wang and Jia Deng · 2020
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing, 2020
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Magnushammer: A transformer-based approach to premise selection
Maciej Mikuła, Szymon Antoniak, Szymon Tworkowski, Albert Qiaochu Jiang, Jin Peng Zhou, Christian Szegedy, Łukasz Kuciński, Piotr Miłoś, and Yuhuai Wu · 2023
Later among the works it cites.
Machine-learned premise selection for Lean
Bartosz Piotrowski, Ramon Fernández Mir, and Edward Ayers · 2023
Later among the works it cites.
Formal mathematics statement curriculum learning
Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, and Ilya Sutskever · 2023
Later among the works it cites.
Passport: Improving automated formal verification with identifiers
Alex Sanchez-Stern, Emily First, Timothy Zhou, Zhanna Kaufman, Yuriy Brun, and Talia Ringer · 2023
Later among the works it cites.
LeanDojo: Theorem proving with retrieval-augmented language models
Kaiyu Yang, Aidan Swope, Alex Gu, Rahul Chalamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan Prenger, and Anima Anandkumar · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
TacticToe: learning to prove with tactics
Thibault Gauthier, Cezary Kaliszyk, Josef Urban, Ramana Kumar, and Michael Norrish · 2021
Cited alongside, same era.
LISA: Language models of ISAbelle proofs
Albert Qiaochu Jiang, Wenda Li, Jesse Michael Han, and Yuhuai Wu · 2021
Cited alongside, same era.
IsarStep: a benchmark for high-level mathematical reasoning
Wenda Li, Lei Yu, Yuhuai Wu, and Lawrence C Paulson · 2021
Cited alongside, same era.
Mathematical reasoning via self-supervised skip-tree training
Markus Norman Rabe, Dennis Lee, Kshitij Bansal, and Christian Szegedy · 2021
Cited alongside, same era.
Completion of the liquid tensor experiment
Mathlib Community · 2022
Cited alongside, same era.
Zarathustra A. Goertzel, Jan Jakubův, Cezary Kaliszyk, Miroslav Olšák, Jelle Piepenbrock, and Josef Urban · 2022
Cited alongside, same era.
Proof artifact co-training for theorem proving with language models
Jesse Michael Han, Jason Rute, Yuhuai Wu, Edward Ayers, and Stanislas Polu · 2022
Cited alongside, same era.
Burak Yetiştiren, Işık Özsoy, Miray Ayerdem, and Eray Tüzün · 2023
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Closest in time.
Machine learning and information theory concepts towards an ai mathematician, 2024
Yoshua Bengio and Nikolay Malkin · 2024
Closest in time.
Internlm2 technical report, 2024
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiaoyi Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang, Penglong Jiao, Zhenjiang Jin, Zhikai Lei, Jiaxing Li, Jingwen Li, Linyang Li, Shuaibin Li, Wei Li, Yining Li, Hongwei Liu, Jiangning Liu, Jiawei Hong, Kaiwen Liu, Kuikun Liu, Xiaoran Liu, Chengqi Lv, Haijun Lv, Kai Lv, Li Ma, Runyuan Ma, Zerun Ma, Wenchang Ning, Linke Ouyang, Jiantao Qiu, Yuan Qu, Fukai Shang, Yunfan Shao, Demin Song, Zifan Song, Zhihao Sui, Peng Sun, Yu Sun, Huanze Tang, Bin Wang, Guoteng Wang, Jiaqi Wang, Jiayu Wang, Rui Wang, Yudong Wang, Ziyi Wang, Xingjian Wei, Qizhen Weng, Fan Wu, Yingtong Xiong, Chao Xu, Ruiliang Xu, Hang Yan, Yirong Yan, Xiaogui Yang, Haochen Ye, Huaiyuan Ying, Jia Yu, Jing Yu, Yuhang Zang, Chuyu Zhang, Li Zhang, Pan Zhang, Peng Zhang, Ruijie Zhang, Shuo Zhang, Songyang Zhang, Wenjian Zhang, Wenwei Zhang, Xingcheng Zhang, Xinyue Zhang, Hui Zhao, Qian Zhao, Xiaomeng Zhao, Fengzhe Zhou, Zaida Zhou, Jingming Zhuo, Yicheng Zou, Xipeng Qiu, Yu Qiao, and Dahua Lin · 2024
Closest in time.
Evaluating language models for mathematics through interactions
Katherine M. Collins, Albert Q. Jiang, Simon Frieder, Lionel Wong, Miri Zilka, Umang Bhatt, Thomas Lukasiewicz, Yuhuai Wu, Joshua B. Tenenbaum, William Hart, Timothy Gowers, Wenda Li, Adrian Weller, and Mateja Jamnik · 2024
Closest in time.
A survey on in-context learning, 2024
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, Baobao Chang, Xu Sun, Lei Li, and Zhifang Sui · 2024
Closest in time.
Data for mathematical copilots: Better ways of presenting proofs for machine learning, 2024
Simon Frieder, Jonas Bayer, Katherine M. Collins, Julius Berner, Jacob Loader, András Juhász, Fabian Ruehle, Sean Welleck, Gabriel Poesia, Ryan-Rhys Griffiths, Adrian Weller, Anirudh Goyal, Thomas Lukasiewicz, and Timothy Gowers · 2024
Closest in time.
Energy efficient convolutions with temporal arithmetic
Rhys Gretsch, Peiyang Song, Advait Madhavan, Jeremy Lau, and Timothy Sherwood · 2024
Closest in time.
Leanagent: Lifelong learning for formal theorem proving, 2024
Adarsh Kumarappan, Mo Tiwari, Peiyang Song, Robert Joseph George, Chaowei Xiao, and Anima Anandkumar · 2024
Closest in time.
A survey on deep learning for theorem proving, 2024
Zhaoyu Li, Jialiang Sun, Logan Murphy, Qidong Su, Zenan Li, Xian Zhang, Kaiyu Yang, and Xujie Si · 2024
Closest in time.
Lean-star: Learning to interleave thinking and proving, 2024
Haohan Lin, Zhiqing Sun, Yiming Yang, and Sean Welleck · 2024
Closest in time.
Magnushammer: A transformer-based approach to premise selection, 2024
Maciej Mikuła, Szymon Antoniak, Szymon Tworkowski, Albert Qiaochu Jiang, Jin Peng Zhou, Christian Szegedy, Łukasz Kuciński, Piotr Miłoś, and Yuhuai Wu · 2024
Closest in time.
OpenAI · 2024
Closest in time.
Learning formal mathematics from intrinsic motivation, 2024
Gabriel Poesia, David Broman, Nick Haber, and Noah D. Goodman · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Google Gemini Team · 2024
Closest in time.
Proving theorems recursively, 2024
Haiming Wang, Huajian Xin, Zhengying Liu, Wenda Li, Yinya Huang, Jianqiao Lu, Zhicheng Yang, Jing Tang, Jian Yin, Zhenguo Li, and Xiaodan Liang · 2024
Closest in time.
Huajian Xin, Z. Z. Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Qihao Zhu, Dejian Yang, Zhibin Gou, Z. F. Wu, Fuli Luo, and Chong Ruan · 2024
Closest in time.
Formal mathematical reasoning: A new frontier in ai, 2024
Kaiyu Yang, Gabriel Poesia, Jingxuan He, Wenda Li, Kristin Lauter, Swarat Chaudhuri, and Dawn Song · 2024
Closest in time.
Lean workbook: A large-scale lean problem set formalized from natural language math problems, 2024
Huaiyuan Ying, Zijian Wu, Yihan Geng, Jiayu Wang, Dahua Lin, and Kai Chen · 2024
Closest in time.
Galore: Memory-efficient llm training by gradient low-rank projection, 2024
Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, and Yuandong Tian · 2024
Closest in time.
Stp: Self-play llm theorem provers with iterative conjecturing and proving, 2025
Kefan Dong and Tengyu Ma · 2025
Closest in time.
Goedel-prover: A frontier model for open-source automated theorem proving, 2025
Yong Lin, Shange Tang, Bohan Lyu, Jiayun Wu, Hongzhou Lin, Kaiyu Yang, Jia Li, Mengzhou Xia, Danqi Chen, Sanjeev Arora, and Chi Jin · 2025
Closest in time.