Fetching the paper…
Reading the bibliography…
We investigate the mathematical capabilities of two iterations of ChatGPT (released 9-January-2023 and 30-January-2023) and of GPT-4 by testing them on publicly available datasets, as well as hand-crafted ones, using a novel methodology.
Some studies in machine learning using the game of checkers
A. L. Samuel · 1959
Earlier work this paper cites.
Principles of Mathematical Analysis
W. Rudin · 1976
Earlier work this paper cites.
Functional analysis
Walter Rudin · 1991
Earlier work this paper cites.
Problem-Solving Strategies
Arthur Engel · 1998
Earlier work this paper cites.
Learning from previous proof experience: A survey
Jörg Denzinger, Matthias Fuchs, Christoph Goller, and Stephan Schulz · 1999
Earlier work this paper cites.
Topology
James R. Munkres · 2000
Earlier work this paper cites.
History of interactive theorem proving
John Harrison, Josef Urban, and Freek Wiedijk · 2014
Earlier work this paper cites.
Linear algebra done right
Sheldon Axler · 2015
Earlier work this paper cites.
Machine-learning the string landscape
Yang-Hui He · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Deep learning for symbolic mathematics
Guillaume Lample and François Charton · 2019
Earlier work this paper cites.
MathQA: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Shanchuan Lin, and Rik Koncel-Kedziorski et al · 2019
Earlier work this paper cites.
Probability: Theory and Examples
Rick Durrett · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, and Jared D Kaplan et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
The Lean mathematical library
The mathlib Community · 2020
Earlier work this paper cites.
Mathematical reasoning via self-supervised skip-tree training
Markus N. Rabe, Dennis Lee, Kshitij Bansal, and Christian Szegedy · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, and Steven Basart et al · 2021
Cited alongside, same era.
Advancing mathematics by guiding human intuition with AI
Alex Davies, Petar Veličković, Lars Buesing, Sam Blackwell, and Daniel Zheng et al · 2021
Cited alongside, same era.
Learning advanced mathematical computations from examples
Francois Charton, Amaury Hayat, and Guillaume Lample · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, and Heewoo Jun et al · 2021
Cited alongside, same era.
Measuring and improving BERT’s mathematical abilities by predicting the order of reasoning
Piotr Piękos, Mateusz Malinowski, and Henryk Michalewski · 2021
Cited alongside, same era.
Das Ende von Google, wie wir es kannten
Sascha Lobo · 2023
Closest in time.
The ChatGPT bot is causing panic now – but it’ll soon be as mundane a tool as Excel
John Naughton · 2023
Closest in time.
I made ChatGPT take a full SAT test. Here’s how it did: [Image attached] [Tweet]. Twitter
teddy [@teddynpc] · 2023
Closest in time.
It’s amusing when ChatGPT makes ridiculous mathematical mistakes. But of course, it’s more interesting to find out what it can do well. Here’s one example that wasn’t bad: I gave it a very rough outline of a proof and asked it to fill in the details [Tweet]. Twitter
Timothy Gowers [@wtgowers] · 2023
Closest in time.
OpenAI (2023) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Naturalproofs: Mathematical theorem proving in natural language
Sean Welleck, Jiacheng Liu, Ronan Le Bras, Hannaneh Hajishirzi, Yejin Choi, and Kyunghyun Cho · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, and Henrique Ponde de Oliveira Pinto et al · 2021
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford · 2021
Cited alongside, same era.
The Brilliance and Weirdness of ChatGPT
Kevon Roose · 2022
Cited alongside, same era.
Performance of ChatGPT on USMLE: Potential for AI-Assisted Medical Education Using Large Language Models
Tiffany H. Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, and Lorie De Leon et al · 2022
Cited alongside, same era.
Machine Learning Class Numbers of Real Quadratic Fields
Malik Amir, Yang-Hui He, Kyu-Hwan Lee, Thomas Oliver, and Eldar Sultanow · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Johan Andreassen, David Dohan, Ethan Dyer, and Henryk Michalewski et al · 2022
Cited alongside, same era.
What is the IQ of ChatGPT?
David Rozado · 2023
Closest in time.
Would Chat GPT3 Get a Wharton MBA? A Prediction Based on Its Performance in the Operations Management Course
Christian Terwiesch · 2023
Closest in time.
ChatGPT – Release Notes
Natalie · 2023
Closest in time.
Does ChatGPT code LaTeX and write proofs? Youtube
Tranquil Sea Of Math · 2023
Closest in time.
Huh. ChatGPT confidently gives the right kind of reasoning to solve this math problem, but whiffs on the algebra in the middle and gets the answer wrong [Tweet]. Twitter
Richard Van Noorden @richvn@mastodon.social [@Richvn] · 2023
Closest in time.
ChatGPT Usage and Limitations
Amos Azaria · 2023
Closest in time.
Mathematics, word problems, common sense, and artificial intelligence
Ernest Davis · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
InstructGPT: Training Language Models to Follow Instructions with Human Feedback
Carroll Wainwright and Ryan Lowe · 2023
Closest in time.
If text-davinci-001 is a rough approximate to the model reported in the NeurIPS 2020 paper, and text-davinci-002 is InstructGPT in the 2022 preprint, then what is just "davinci"? Trying to reproduce results from a time before this naming existed [Tweet]. Twitter
Sarah Wiegreffe (sigmoid.social/@sarah) [@sarahwiegreffe] · 2023
Closest in time.
GPT-4 API waitlist
OpenAI · 2023
Closest in time.
Documentation - Models
OpenAI · 2023
Closest in time.
OpenAI API Reference - Chat Completion Endpoint
OpenAI · 2023
Closest in time.