Fetching the paper…
Reading the bibliography…
Large language models (LLMs) such as ChatGPT have received immense interest for their general-purpose language understanding and, in particular, their ability to generate high-quality text or computer code.
Functional analysis
W. Rudin · 1991
Earlier work this paper cites.
The Coq proof assistant reference manual: Version 6.1
B. Barras, S. Boutin, C. Cornes, J. Courant, J.-C. Filliatre, E. Gimenez, H. Herbelin, G. Huet, C. Munoz, C. Murthy, et al · 1997
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, and P. Vincent · 2000
Earlier work this paper cites.
Topology
J. R. Munkres · 2000
Earlier work this paper cites.
The Lean theorem prover (system description)
L. de Moura, S. Kong, J. Avigad, F. Van Doorn, and J. von Raumer · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
R. Sennrich, B. Haddow, and A. Birch · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Probability: Theory and Examples
R. Durrett · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Cited alongside, same era.
Energy and policy considerations for deep learning in NLP
E. Strubell, A. Ganesh, and A. McCallum · 2019
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline
I. Tenney, D. Das, and E. Pavlick · 2019
Cited alongside, same era.
Learning to prove theorems via interacting with proof assistants
K. Yang and J. Deng · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Evaluating language models for mathematics through interactions
K. M. Collins, A. Q. Jiang, S. Frieder, L. Wong, M. Zilka, U. Bhatt, T. Lukasiewicz, Y. Wu, J. B. Tenenbaum, W. Hart, et al · 2023
Closest in time.
GPTs are GPTs: An early look at the labor market impact potential of large language models
T. Eloundou, S. Manning, P. Mishkin, and D. Rock · 2023
Closest in time.
Baldur: whole-proof generation and repair with large language models
E. First, M. N. Rabe, T. Ringer, and Y. Brun · 2023
Closest in time.
LLM vs ITP
S. Frieder, M. Alawadhi, Trimmel, Rashid, and K. Gy · 2023
Closest in time.
Mathematical capabilities of ChatGPT
S. Frieder, L. Pinchetti, R.-R. Griffiths, T. Salvatori, T. Lukasiewicz, P. C. Petersen, A. Chevalier, and J. Berner · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The Pile: An 800GB dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Cited alongside, same era.
OpenAI’s GPT-3 language model: A technical overview, 2020
C. Li · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Cited alongside, same era.
Measuring mathematical problem solving with the MATH dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Multi-head or single-head? an empirical comparison for transformer training
L. Liu, J. Liu, and J. Han · 2021
Cited alongside, same era.
GLaM: Efficient scaling of language models with mixture-of-experts
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, et al · 2022
Cited alongside, same era.
Transformer language models without positional encodings still learn positional information
A. Haviv, O. Ram, O. Press, P. Izsak, and O. Levy · 2022
Cited alongside, same era.
Humans are still better than ChatGPT: Case of the IEEEXtreme competition
A. Koubaa, B. Qureshi, A. Ammar, Z. Khan, W. Boulila, and L. Ghouti · 2023
Closest in time.
Recent advances in natural language processing via large pre-trained language models: A survey
B. Min, H. Ross, E. Sulem, A. P. B. Veyseh, T. H. Nguyen, O. Sainz, E. Agirre, I. Heintz, and D. Roth · 2023
Closest in time.
GPT-4 technical report
OpenAI · 2023
Closest in time.
PanGu- Σ \Sigma : Towards trillion parameter language model with sparse heterogeneous computing
X. Ren, P. Zhou, X. Meng, X. Huang, Y. Wang, W. Wang, P. Li, X. Zhang, A. Podolskiy, G. Arshinov, et al · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
Is deep learning a useful tool for the pure mathematician?
G. Williamson · 2023
Closest in time.
LeanDojo: Theorem proving with retrieval-augmented language models
K. Yang, A. M. Swope, A. Gu, R. Chalamala, P. Song, S. Yu, S. Godil, R. Prenger, and A. Anandkumar · 2023
Closest in time.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Closest in time.