Fetching the paper…
Reading the bibliography…
We propose the Quantization Model of neural scaling laws, explaining both the observed power law dropoff of loss with model and data size, and also the sudden emergence of new capabilities with scale.
“Society of mind”
Marvin Minsky · 1988
Earlier work this paper cites.
Thao Nguyen, Maithra Raghu and Simon Kornblith · 2010
Earlier work this paper cites.
“Phase transitions in machine learning”
Lorenza Saitta, Attilio Giordana and Antoine Cornuejols · 2011
Earlier work this paper cites.
“Scikit-learn: Machine Learning in Python”
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot and E. Duchesnay · 2011
Earlier work this paper cites.
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson and John Hopcroft · 2016
Earlier work this paper cites.
“Deep learning scaling is predictable, empirically”
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang and Yanqi Zhou · 2017
Earlier work this paper cites.
“The lottery ticket hypothesis: Finding sparse, trainable neural networks”
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
“A constructive prediction of the generalization error across scales”
Jonathan Rosenfeld, Amir Rosenfeld, Yonatan Belinkov and Nir Shavit · 2019
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu and Dario Amodei · 2020
Earlier work this paper cites.
“Scaling laws for autoregressive generative modeling”
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom Brown, Prafulla Dhariwal and Scott Gray · 2020
Earlier work this paper cites.
“Zoom In: An Introduction to Circuits” https://distill.pub/2020/circuits/zoom-in
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov and Shan Carter · 2020
Earlier work this paper cites.
“Curve detectors”
Nick Cammarata, Gabriel Goh, Shan Carter, Ludwig Schubert, Michael Petrov and Chris Olah · 2020
Earlier work this paper cites.
“The pile: An 800gb dataset of diverse text for language modeling”
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite and Noa Nabeshima · 2020
Earlier work this paper cites.
“Spectrum dependent learning curves in kernel regression and wide neural networks”
Blake Bordelon, Abdulkadir Canatar and Cengiz Pehlevan · 2020
Earlier work this paper cites.
“Data and parameter scaling laws for neural machine translation”
Mitchell Gordon, Kevin Duh and Jared Kaplan · 2021
Earlier work this paper cites.
“A Mathematical Framework for Transformer Circuits” https://transformer-circuits.pub/2021/framework/index.html
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish and Chris Olah · 2021
Earlier work this paper cites.
“The scaling hypothesis”, 2021
Gwern Branwen · 2021
Earlier work this paper cites.
“Explaining neural scaling laws”
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee and Utkarsh Sharma · 2021
Cited alongside, same era.
Marcus Hutter · 2021
Cited alongside, same era.
“Scaling laws for acoustic models”
Jasha Droppo and Oguz Elibol · 2021
Cited alongside, same era.
“Scaling vision transformers”
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby and Lucas Beyer · 2022
Cited alongside, same era.
“Training compute-optimal large language models”
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego Casas, Lisa Hendricks, Johannes Welbl and Aidan Clark · 2022
“Beyond the imitation game: Quantifying and extrapolating the capabilities of language models”
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Shoeb, Abubakar Abid, Adam Fisch, Adam Brown, Adam Santoro, Aditya Gupta and Adrià Garriga-Alonso · 2022
Later among the works it cites.
“Data distributional properties drive emergent in-context learning in transformers”
Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland and Felix Hill · 2022
Later among the works it cites.
“Understanding Scaling Laws for Recommendation Models”
Newsha Ardalani, Carole-Jean Wu, Zeliang Chen, Bhargav Bhushanam and Adnan Aziz · 2022
Later among the works it cites.
“Progress measures for grokking via mechanistic interpretability”
Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith and Jacob Steinhardt · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Emergent Abilities of Large Language Models” Survey Certification
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean and William Fedus · 2022
Cited alongside, same era.
“Future ML Systems Will Be Qualitatively Different” Accessed: 2023-10-26, 2022
Jacob Steinhardt · 2022
Cited alongside, same era.
“Predictability and surprise in large generative models”
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain and Nelson Elhage · 2022
Cited alongside, same era.
“Emergent world representations: Exploring a sequence model trained on a synthetic task”
Kenneth Li, Aspen Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister and Martin Wattenberg · 2022
Cited alongside, same era.
“Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small”
Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris and Jacob Steinhardt · 2022
Cited alongside, same era.
“Graphical clusterability and local specialization in deep neural networks”
Stephen Casper, Shlomi Hod, Daniel Filan, Cody Wild, Andrew Critch and Stuart Russell · 2022
Cited alongside, same era.
“In-context Learning and Induction Heads” https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish and Chris Olah · 2022
Cited alongside, same era.
Tom Lieberum, Matthew Rahtz, János Kramár, Geoffrey Irving, Rohin Shah and Vladimir Mikulik · 2023
Closest in time.
“Towards Monosemanticity: Decomposing Language Models With Dictionary Learning” https://transformer-circuits.pub/2023/monosemantic-features/index.html
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah Burke, Tristan Hume, Shan Carter, Tom Henighan and Christopher Olah · 2023
Closest in time.
“Discovering Knowledge-Critical Subnetworks in Pretrained Language Models”
Deniz Bayazit, Negar Foroutan, Zeming Chen, Gail Weiss and Antoine Bosselut · 2023
Closest in time.
“Rosetta Neurons: Mining the Common Units in a Model Zoo” arXiv:2306.09346 [cs]
Amil Dravid, Yossi Gandelsman, Alexei. Efros and Assaf Shocher · 2023
Closest in time.
“Pythia: A suite for analyzing large language models across training and scaling”
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Khan, Shivanshu Purohit, USVSN Prashanth and Edward Raff · 2023
Closest in time.
“Precision Machine Learning”
Eric. Michaud, Ziming Liu and Max Tegmark · 2023
Closest in time.
“Are Emergent Abilities of Large Language Models a Mirage?”
Rylan Schaeffer, Brando Miranda and Sanmi Koyejo · 2023
Closest in time.
“A theory for emergence of complex skills in language models”
Sanjeev Arora and Anirudh Goyal · 2023
Closest in time.
“On the stepwise nature of self-supervised learning”
James Simon, Maksis Knutins, Liu Ziyin, Daniel Geisz, Abraham Fetterman and Joshua Albrecht · 2023
Closest in time.
“Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models”
Mayee Chen, Nicholas Roberts, Kush Bhatia, Jue Wang, Ce Zhang, Frederic Sala and Christopher Ré · 2023
Closest in time.
“TinyStories: How Small Can Language Models Be and Still Speak Coherent English?”
Ronen Eldan and Yuanzhi Li · 2023
Closest in time.
“Scaling Laws Literature Review” Accessed: 2023-01-31, https://epochai.org/blog/scaling-laws-literature-review , 2023
Pablo Villalobos · 2023
Closest in time.