Fetching the paper…
Reading the bibliography…
Despite the success of large language models (LLMs), the task of theorem proving still remains one of the hardest reasoning tasks that is far from being fully solved.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Kevin Buzzard, Johan Commelin, and Patrick Massot · 1910
Earlier work this paper cites.
Deep Learning for Symbolic Mathematics
Guillaume Lample and François Charton · 1912
Earlier work this paper cites.
Rings of sets
Garrett Birkhoff · 1937
Earlier work this paper cites.
The mind of mechanical man
Geoffrey Jefferson · 1949
Earlier work this paper cites.
Three models for the description of language
Noam Chomsky · 1956
Earlier work this paper cites.
A machine program for theorem-proving
Martin Davis, George Logemann, and Donald Loveland · 1962
Earlier work this paper cites.
A Machine-Oriented Logic Based on the Resolution Principle
J. A. Robinson · 1965
Earlier work this paper cites.
The Formulae-as-Types Notion of Construction
William Alvin Howard · 1980
Earlier work this paper cites.
The calculus of constructions
Thierry Coquand and Gérard Huet · 1986
Earlier work this paper cites.
The calculus of constructions
Thierry Coquand and Gérard Huet · 1988
Earlier work this paper cites.
Inductively defined types
Thierry Coquand and Christine Paulin · 1990
Earlier work this paper cites.
Isabelle a Generic Theorem Prover
Lawrence C. Paulson · 1994
Earlier work this paper cites.
The Coq Proof Assistant Reference Manual : Version 6.1
Bruno Barras, Samuel Boutin, Cristina Cornes, Judicaël Courant, Jean-Christophe Filliâtre, Eduardo Giménez, Hugo Herbelin, Gérard Huet, César Muñoz, Chetan Murthy, Catherine Parent, Christine Paulin-Mohring, Amokrane Saïbi, and Benjamin Werner · 1997
Earlier work this paper cites.
The design and implementation of vampire
Alexandre Riazanov and Andrei Voronkov · 2002
Earlier work this paper cites.
E - a brainiac theorem prover
Stephan Schulz · 2002
Earlier work this paper cites.
A proposal for the dartmouth summer research project on artificial intelligence, august 31, 1955
John McCarthy, Marvin L Minsky, Nathaniel Rochester, and Claude E Shannon · 2006
Earlier work this paper cites.
INT: An Inequality Benchmark for Evaluating Generalization in Theorem Proving
Yuhuai Wu, Albert Qiaochu Jiang, Jimmy Ba, and Roger Grosse · 2007
Earlier work this paper cites.
Z3: An efficient smt solver
Leonardo De Moura and Nikolaj Bjørner · 2008
Earlier work this paper cites.
Z3: An Efficient SMT Solver
Leonardo de Moura and Nikolaj Bjørner · 2008
Earlier work this paper cites.
HOL light: An overview
John Harrison · 2009
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever · 2009
Earlier work this paper cites.
Generative Language Modeling for Automated Theorem Proving
Stanislas Polu and Ilya Sutskever · 2009
Earlier work this paper cites.
Spass version 3.5
Christoph Weidenbach, Dilyana Dimova, Arnaud Fietzke, Rohit Kumar, Martin Suda, and Patrick Wischnewski · 2009
Earlier work this paper cites.
Licensing the mizar mathematical library
Jesse Alama, Michael Kohlhase, Lionel Mamane, Adam Naumowicz, Piotr Rudnicki, and Josef Urban · 2011
Earlier work this paper cites.
First-Order Theorem Proving and Vampire
Laura Kovács and Andrei Voronkov · 2013
Earlier work this paper cites.
The Lean Theorem Prover (System Description)
Leonardo de Moura, Soonho Kong, Jeremy Avigad, Floris van Doorn, and Jakob von Raumer · 2015
Earlier work this paper cites.
Introduction to Lattice Theory with Computer Science Applications
Vijay K. Garg · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
DeepMath - deep sequence models for premise selection
Alexander A. Alemi, François Chollet, Niklas Een, Geoffrey Irving, Christian Szegedy, and Josef Urban · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
A metaprogramming framework for formal verification
Gabriel Ebner, Sebastian Ullrich, Jared Roesch, Jeremy Avigad, and Leonardo de Moura · 2017
Cited alongside, same era.
Tactictoe: Learning to reason with HOL4 tactics
Thibault Gauthier, Cezary Kaliszyk, and Josef Urban · 2017
Cited alongside, same era.
Deep network guided proof search
Sarah M. Loos, Geoffrey Irving, Christian Szegedy, and Cezary Kaliszyk · 2017
Cited alongside, same era.
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Recurrent memory transformer, 2022
Aydar Bulatov, Yuri Kuratov, and Mikhail S. Burtsev · 2022
Later among the works it cites.
Discovering the Hidden Vocabulary of DALLE-2
Giannis Daras and Alexandros G. Dimakis · 2022
Later among the works it cites.
The independence of the continuum hypothesis in isabelle/zf
Emmanuel Gunther, Miguel Pagano, Pedro Sánchez Terraf, and Matías Steinberg · 2022
Later among the works it cites.
Proof artifact co-training for theorem proving with language models
Jesse Michael Han, Jason Rute, Yuhuai Wu, Edward W. Ayers, and Stanislas Polu · 2022
Later among the works it cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
First experiments with neural translation of informal to formal mathematics
Qingxiang Wang, Cezary Kaliszyk, and Josef Urban · 2018
Cited alongside, same era.
Holist: An environment for machine learning of higher order logic theorem proving
Kshitij Bansal, Sarah M. Loos, Markus N. Rabe, Christian Szegedy, and Stewart Wilcox · 2019
Cited alongside, same era.
Formalizing the solution to the cap set problem
Sander R. Dahmen, Johannes Hölzl, and Robert Y. Lewis · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension, 2019
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2019
Cited alongside, same era.
Metamath: a computer language for mathematical proofs
Norman Megill and David A Wheeler · 2019
Cited alongside, same era.
Learning to prove theorems via interacting with proof assistants
Kaiyu Yang and Jia Deng · 2019
Cited alongside, same era.
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Later among the works it cites.
HyperTree Proof Search for Neural Theorem Proving
Guillaume Lample, Marie-Anne Lachaux, Thibaut Lavril, Xavier Martinet, Amaury Hayat, Gabriel Ebner, Aurélien Rodriguez, and Timothée Lacroix · 2022
Later among the works it cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V. Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra · 2022
Later among the works it cites.
Competition-Level Code Generation with AlphaCode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals · 2022
Later among the works it cites.
Formal Mathematics Statement Curriculum Learning
Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, and Ilya Sutskever · 2022
Later among the works it cites.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas · 2022
Later among the works it cites.
Autoformalization with large language models
Yuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus Rabe, Charles Staats, Mateja Jamnik, and Christian Szegedy · 2022
Later among the works it cites.
Expression syntax information bottleneck for math word problems
Jing Xiong, Chengming Li, Min Yang, Xiping Hu, and Bin Hu · 2022
Later among the works it cites.
Mathematics and the formal turn
Jeremy Avigad · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Closest in time.
Scaling transformer to 1m tokens and beyond with rmt, 2023
Aydar Bulatov, Yuri Kuratov, and Mikhail S. Burtsev · 2023
Closest in time.
Large language models as tool makers
Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou · 2023
Closest in time.
Baldur: Whole-proof generation and repair with large language models
Emily First, Markus N. Rabe, Talia Ringer, and Yuriy Brun · 2023
Closest in time.
Fimo: A challenge formal dataset for automated theorem proving
Chengwu Liu, Jianhao Shen, Huajian Xin, Zhengying Liu, Ye Yuan, Haiming Wang, Wei Ju, Chuanyang Zheng, Yichun Yin, Lin Li, et al · 2023
Closest in time.
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang · 2023
Closest in time.
Magnushammer: A transformer-based approach to premise selection
Maciej Mikuła, Szymon Antoniak, Szymon Tworkowski, Albert Qiaochu Jiang, Jin Peng Zhou, Christian Szegedy, Łukasz Kuciński, Piotr Miłoś, and Yuhuai Wu · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior, 2023
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2023
Closest in time.
Large ai models in health informatics: Applications, challenges, and the future
Jianing Qiu, Lin Li, Jiankai Sun, Jiachuan Peng, Peilun Shi, Ruiyang Zhang, Yinzhao Dong, Kyle Lam, Frank P-W Lo, Bo Xiao, et al · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face, 2023
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang · 2023
Closest in time.
Recursion (computer science) — Wikipedia, the free encyclopedia, 2023
Wikipedia contributors · 2023
Closest in time.
Leandojo: Theorem proving with retrieval-augmented language models
Kaiyu Yang, Aidan M Swope, Alex Gu, Rahul Chalamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan Prenger, and Anima Anandkumar · 2023
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models, 2023
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu · 2023
Closest in time.
Decomposing the enigma: Subgoal-based demonstration learning for formal theorem proving
Xueliang Zhao, Wenda Li, and Lingpeng Kong · 2023
Closest in time.
Progressive-hint prompting improves reasoning in large language models
Chuanyang Zheng, Zhengying Liu, Enze Xie, Zhenguo Li, and Yu Li · 2023
Closest in time.