Fetching the paper…
Reading the bibliography…
A major driver of AI products today is the fact that new skills emerge in language models when their parameter set and training corpora are scaled up.
Syntactic Structures
N Chomsky · 1957
Earlier work this paper cites.
Procedures as a representation for data in a computer program for understanding natural language
Terry Winograd · 1971
Earlier work this paper cites.
Logic and conversation
H. P. Grice · 1975
Earlier work this paper cites.
Language and the problems of knowledge
N Chomsky · 1988
Earlier work this paper cites.
Learning curves: Asymptotic values and rate of convergence
Corinna Cortes, Lawrence D Jackel, Sara Solla, Vladimir Vapnik, and John Denker · 1993
Earlier work this paper cites.
Framing in Discourse
D Tannen (ed) · 1994
Earlier work this paper cites.
Surface Structure and Interpretation
Mark Steedman · 1996
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
Mathematical foundations for a compositional distributional model of meaning
B Coecke, M Sadrzadeh, and S Clark · 2010
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
The Probabilistic Method (4th Ed)
N Alon and J Spencer · 2016
Cited alongside, same era.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Cited alongside, same era.
Cambridge Handbook of Morphology
A Hippisley and G Stump · 2017
Cited alongside, same era.
Language Assessment: Principles and Classroom Practice (6th Ed)
HD Brown · 2018
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Predictability and surprise in large generative models
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily Bender, Timnit Gebru, Andrea McMillan-Major, and Shmargaret Shmitchel · 2021
Cited alongside, same era.
A mathematical exploration of why language models help solve downstream tasks
Nikunj Saunshi, Sadhika Malladi, and Sanjeev Arora · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Cited alongside, same era.
Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, and Ves Stoyanov · 2022
Later among the works it cites.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Closest in time.
Scaling data-constrained language models, 2023
Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Are emergent abilities of large language models a mirage?
Ryan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Closest in time.