Fetching the paper…
Reading the bibliography…
There is much excitement about the opportunity to harness the power of large language models (LLMs) when building problem-solving assistants.
The logic theory machine–a complex information processing system
A. Newell and H. Simon · 1956
Earlier work this paper cites.
Toward mechanical mathematics
H. Wang · 1960
Earlier work this paper cites.
A machine program for theorem-proving
M. Davis, G. Logemann, and D. Loveland · 1962
Earlier work this paper cites.
Truth and proof
A. Tarski · 1969
Earlier work this paper cites.
Non-resolution theorem proving
W. W. Bledsoe · 1977
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, O. Bousquet, Q. Le, and E. Chi · 1977
Earlier work this paper cites.
The Computer Modelling of Mathematical Reasoning
A. Bundy · 1983
Earlier work this paper cites.
The use of explicit plans to guide inductive proofs
A. Bundy · 1988
Earlier work this paper cites.
Mind children: The future of robot and human intelligence
H. Moravec · 1988
Earlier work this paper cites.
Rippling: A heuristic for guiding inductive proofs
A. Bundy, A. Stevens, F. V. Harmelen, A. Ireland, and A. Smaill · 1993
Earlier work this paper cites.
Implementing tactics and tacticals in a higher-order logic programming language
A. P. Felty · 1993
Earlier work this paper cites.
Interactive theorem proving: An empirical study of user activity
J. S. Aitken, P. Gray, T. Melham, and M. Thomas · 1998
Earlier work this paper cites.
A tactic language for the system coq
D. Delahaye · 2000
Earlier work this paper cites.
Evaluating general purpose automated theorem proving systems
G. Sutcliffe and C. Suttner · 2001
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. Zhu · 2002
Earlier work this paper cites.
E - a brainiac theorem prover
S. Schulz · 2002
Earlier work this paper cites.
Automation bias in intelligent time critical decision support systems
M. L. Cummings · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
A proof of the kepler conjecture
T. C. Hales · 2005
Earlier work this paper cites.
Z3: An efficient smt solver
L. M. de Moura and N. S. Bjørner · 2008
Earlier work this paper cites.
Generative language modeling for automated theorem proving
S. Polu and I. Sutskever · 2009
Earlier work this paper cites.
Three years of experience with sledgehammer, a practical link between automatic and interactive theorem provers
L. C. Paulson · 2010
Earlier work this paper cites.
A fully automatic problem solver with human-style output
M. Ganesalingam and W. T. Gowers · 2013
Earlier work this paper cites.
The flyspeck project, Aug 2014
T. Hales · 2014
Earlier work this paper cites.
History of interactive theorem proving
J. Harrison, J. Urban, and F. Wiedijk · 2014
Earlier work this paper cites.
A usability evaluation of interactive theorem provers using focus groups
B. Beckert, S. Grebing, and F. Böhl · 2015
Earlier work this paper cites.
Use and misuse of the likert item responses and other ordinal measures
P. A. Bishop and R. L. Herron · 2015
Earlier work this paper cites.
Communicating crisis uncertainty: A review of the knowledge gaps
B. F. Liu, L. Bartz, and N. Duke · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Risk and Uncertainty Communication
D. Spiegelhalter · 2017
Earlier work this paper cites.
Creative writing with a machine in the loop: Case studies on slogans and stories
E. Clark, A. S. Ross, C. Tan, Y. Ji, and N. A. Smith · 2018
Earlier work this paper cites.
In pursuit of error: A survey of uncertainty visualization evaluation
J. Hullman, X. Qiao, M. Correll, A. Kale, and M. Kay · 2018
Earlier work this paper cites.
Gradio: Hassle-Free Sharing and Testing of ML Models in the Wild
A. Abid, A. Abdalla, A. Abid, D. Khan, A. Alfozan, and J. Zou · 2019
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi · 2019
Earlier work this paper cites.
Will you accept an imperfect ai? exploring designs for adjusting end-user expectations of ai systems
R. Kocielnik, S. Amershi, and P. N. Bennett · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Explainable machine learning in deployment
U. Bhatt, A. Xiang, S. Sharma, A. Weller, A. Taly, Y. Jia, J. Ghosh, R. Puri, J. M. Moura, and P. Eckersley · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Replica: REPL instrumentation for coq analysis
T. Ringer, A. Sanchez-Stern, D. Grossman, and S. Lerner · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Naturalprover: Grounded mathematical proof generation with language models
S. Welleck, J. Liu, X. Lu, H. Hajishirzi, and Y. Choi · 2022
Later among the works it cites.
Autoformalization with large language models
Y. Wu, A. Q. Jiang, W. Li, M. N. Rabe, C. Staats, M. Jamnik, and C. Szegedy · 2022
Later among the works it cites.
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
Y. Xiao, P. P. Liang, U. Bhatt, W. Neiswanger, R. Salakhutdinov, and L.-P. Morency · 2022
Later among the works it cites.
How transparency modulates trust in artificial intelligence
J. Zerilli, U. Bhatt, and A. Weller · 2022
Later among the works it cites.
miniF2F: a cross-system benchmark for formal Olympiad-level mathematics
K. Zheng, J. M. Han, and S. Polu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Cited alongside, same era.
Uncertainty as a form of transparency: Measuring, communicating, and using uncertainty
U. Bhatt, J. Antorán, Y. Zhang, Q. V. Liao, P. Sattigeri, R. Fogliato, G. Melançon, R. Krishnan, J. Stanley, O. Tickoo, et al · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
Advancing mathematics by guiding human intuition with ai
A. Davies, P. Veličković, L. Buesing, S. Blackwell, D. Zheng, N. Tomašev, R. Tanburn, P. Battaglia, C. Blundell, A. Juhász, et al · 2021
Cited alongside, same era.
Github copilot · your ai pair programmer, 2021
Github · 2021
Cited alongside, same era.
A. F. Akyürek, E. Akyürek, A. Madaan, A. Kalyan, P. Clark, D. Wijaya, and N. Tandon · 2023
Closest in time.
PaLM 2 Technical Report, 2023
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier-Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. Abrego, J. Ahn, J. Austin, P. Barham, J. Botha, J. Bradbury, S. Brahma, K. Brooks, M. Catasta, Y. Cheng, C. Cherry, C. A. Choquette-Choo, A. Chowdhery, C. Crepy, S. Dave, M. Dehghani, S. Dev, J. Devlin, M. Díaz, N. Du, E. Dyer, V. Feinberg, F. Feng, V. Fienber, M. Freitag, X. Garcia, S. Gehrmann, L. Gonzalez, G. Gur-Ari, S. Hand, H. Hashemi, L. Hou, J. Howland, A. Hu, J. Hui, J. Hurwitz, M. Isard, A. Ittycheriah, M. Jagielski, W. Jia, K. Kenealy, M. Krikun, S. Kudugunta, C. Lan, K. Lee, B. Lee, E. Li, M. Li, W. Li, Y. Li, J. Li, H. Lim, H. Lin, Z. Liu, F. Liu, M. Maggioni, A. Mahendru, J. Maynez, V. Misra, M. Moussalem, Z. Nado, J. Nham, E. Ni, A. Nystrom, A. Parrish, M. Pellat, M. Polacek, A. Polozov, R. Pope, S. Qiao, E. Reif, B. Richter, P. Riley, A. C. Ros, A. Roy, B. Saeta, R. Samuel, R. Shelby, A. Slone, D. Smilkov, D. R. So, D. Sohn, S. Tokumine, D. Valter, V. Vasudevan, K. Vodrahalli, X. Wang, P. Wang, Z. Wang, T. Wang, J. Wieting, Y. Wu, K. Xu, Y. Xu, L. Xue, P. Yin, J. Yu, Q. Zhang, S. Zheng, C. Zheng, W. Zhou, D. Zhou, S. Petrov, and Y. Wu · 2023
Closest in time.
Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum
J. W. Ayers, A. Poliak, M. Dredze, E. C. Leas, Z. Zhu, J. B. Kelley, D. J. Faix, A. M. Goodman, C. A. Longhurst, M. Hogarth, and D. M. Smith · 2023
Closest in time.
Large language models and the perils of their hallucinations
R. Azamfirei, S. R. Kudchadkar, and J. Fackler · 2023
Closest in time.
Csc413/2516 winter 2023 university of toronto, assignment 1, 2023
J. Ba and B. Wang · 2023
Closest in time.
Learning personalized decision support policies
U. Bhatt, V. Chen, K. M. Collins, P. Kamalaruban, E. Kallina, A. Weller, and A. Talwalkar · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang · 2023
Closest in time.
Rethink reporting of evaluation results in ai
R. Burnell, W. Schellaert, J. Burden, T. D. Ullman, F. Martinez-Plumed, J. B. Tenenbaum, D. Rutar, L. G. Cheke, J. Sohl-Dickstein, M. Mitchell, D. Kiela, M. Shanahan, E. M. Voorhees, A. G. Cohn, J. Z. Leibo, and J. Hernandez-Orallo · 2023
Closest in time.
Open problems and fundamental limitations of reinforcement learning from human feedback
S. Casper, X. Davies, C. Shi, T. Krendl Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, et al · 2023
Closest in time.
A. G. Cohn and J. Hernandez-Orallo · 2023
Closest in time.
Faith and fate: Limits of transformers on compositionality
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jian, B. Y. Lin, P. West, C. Bhagavatula, R. L. Bras, J. D. Hwang, et al · 2023
Closest in time.
Baldur: Whole-proof generation and repair with large language models
E. First, M. N. Rabe, T. Ringer, and Y. Brun · 2023
Closest in time.
Mathematical Capabilities of ChatGPT, 2023
S. Frieder, L. Pinchetti, R.-R. Griffiths, T. Salvatori, T. Lukasiewicz, P. C. Petersen, A. Chevalier, and J. Berner · 2023
Closest in time.
Capturing humans’ mental models of ai: An item response theory approach
M. Kelly, A. Kumar, P. Smyth, and M. Steyvers · 2023
Closest in time.
Openassistant conversations – democratizing large language model alignment, 2023
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z.-R. Tam, K. Stevens, A. Barhoum, N. M. Duc, O. Stanley, R. Nagyfi, S. ES, S. Suri, D. Glushkov, A. Dantuluri, A. Maguire, C. Schuhmann, H. Nguyen, and A. Mattick · 2023
Closest in time.
Causal reasoning and large language models: Opening a new frontier for causality, 2023
E. Kıcıman, R. Ness, A. Sharma, and C. Tan · 2023
Closest in time.
Lampp: Language models as probabilistic priors for perception and action
B. Z. Li, W. Chen, P. Sharma, and J. Andreas · 2023
Closest in time.
Evaluating statistical language models as pragmatic reasoners
B. Lipkin, L. Wong, G. Grand, and J. B. Tenenbaum · 2023
Closest in time.
Magnushammer: A transformer-based approach to premise selection
M. Mikula, S. Antoniak, S. Tworkowski, A. Q. Jiang, J. P. Zhou, C. Szegedy, L. Kucinski, P. Milos, and Y. Wu · 2023
Closest in time.
Explainable ai is dead, long live explainable ai! hypothesis-driven decision support, 2023
T. Miller · 2023
Closest in time.
Co-writing screenplays and theatre scripts with language models: Evaluation by industry professionals
P. Mirowski, K. W. Mathewson, J. Pittman, and R. Evans · 2023
Closest in time.
Parachute: Evaluating interactive human-lm co-writing systems
H. Shen and T. Wu · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Generation probabilities are not enough: Exploring the effectiveness of uncertainty highlighting in ai-powered code completions
H. Vasconcelos, G. Bansal, A. Fourney, Q. V. Liao, and J. Wortman Vaughan · 2023
Closest in time.
L. Wong, G. Grand, A. K. Lew, N. D. Goodman, V. K. Mansinghka, J. Andreas, and J. B. Tenenbaum · 2023
Closest in time.
Interpretability at scale: Identifying causal mechanisms in alpaca
Z. Wu, A. Geiger, C. Potts, and N. D. Goodman · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models, 2023
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan · 2023
Closest in time.