Fetching the paper…
Reading the bibliography…
Increase in data, size, or compute can lead to sudden learning of specific capabilities by a neural network -- a phenomenon often called "emergence''.
Three models for the description of language
Noam Chomsky · 1956
Earlier work this paper cites.
More is different: Broken symmetry and the nature of the hierarchical structure of science
Philip W Anderson · 1972
Earlier work this paper cites.
Learnability and the Vapnik-Chervonenkis dimension
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth · 1989
Earlier work this paper cites.
Statistical mechanics of learning from examples
Hyunjune Sebastian Seung, Haim Sompolinsky, and Naftali Tishby · 1992
Earlier work this paper cites.
A universal theorem on learning curves
Shun-Ichi Amari · 1993
Earlier work this paper cites.
The statistical mechanics of learning a rule
Timothy LH Watkin, Albrecht Rau, and Michael Biehl · 1993
Earlier work this paper cites.
Rigorous learning curve bounds from statistical mechanics
David Haussler, H Sebastian Seung, Michael Kearns, and Naftali Tishby · 1994
Earlier work this paper cites.
Introduction to the theory of computation
Michael Sipser · 1996
Earlier work this paper cites.
Random graphs with arbitrary degree distributions and their applications
Mark EJ Newman, Steven H Strogatz, and Duncan J Watts · 2001
Earlier work this paper cites.
Percolation critical exponents in scale-free networks
Reuven Cohen, Daniel Ben-Avraham, and Shlomo Havlin · 2002
Earlier work this paper cites.
The Structure and Function of Complex Networks
M. E. J. Newman · 2003
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper · 2009
Earlier work this paper cites.
Asymptotic analysis of the stochastic block model for modular networks and its algorithmic applications
Aurelien Decelle, Florent Krzakala, Cristopher Moore, and Lenka Zdeborová · 2011
Earlier work this paper cites.
Probabilistic context-free grammars (pcfgs)
Michael Collins · 2013
Earlier work this paper cites.
Spectral thresholds in the bipartite stochastic block model
Laura Florescu and Will Perkins · 2016
Earlier work this paper cites.
Community detection and stochastic block models: recent developments
Emmanuel Abbe · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 2020
Earlier work this paper cites.
A theory of universal learning
Olivier Bousquet, Steve Hanneke, Shay Moran, Ramon Van Handel, and Amir Yehudayoff · 2021
Earlier work this paper cites.
Self-supervised learning of deep visual representations
Mathilde Caron · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Earlier work this paper cites.
A Mathematical Framework for Transformer Circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, …, and Chris Olah · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Earlier work this paper cites.
Do As I Can and Not As I Say: Grounding Language in Robotic Affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, …, and Andy Zeng · 2022
Earlier work this paper cites.
Hidden progress in deep learning: SGD learns parities near the computational limit
Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Earlier work this paper cites.
On the Opportunities and Risks of Foundation Models, jul 2022
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, …, and Percy Liang · 2022
Earlier work this paper cites.
Predictability and surprise in large generative models
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al · 2022
Earlier work this paper cites.
An empirical analysis of compute-optimal large language model training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Earlier work this paper cites.
General-purpose in-context learning by meta-learning transformers
Louis Kirsch, James Harrison, Jascha Sohl-Dickstein, and Luke Metz · 2022
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2022
Cited alongside, same era.
In-context Learning and Induction Heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, …, and Chris Olah · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Regulating the Risks of AI
Margot Kaminski · 2023
Later among the works it cites.
Are Emergent Abilities in Large Language Models just In-Context Learning?
Sheng Lu, Irina Bigoulaeva, Rachneet Sachdeva, Harish Tayyar Madabushi, and Iryna Gurevych · 2023
Later among the works it cites.
Mind your language (model): Fact-checking llms and their role in nlp research and practice
Alexandra Sasha Luccioni and Anna Rogers · 2023
Later among the works it cites.
A Tale of Two Circuits: Grokking as Competition of Sparse and Dense Subnetworks
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Later among the works it cites.
The quantization model of neural scaling
Eric J Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Cited alongside, same era.
Benchmarking compositionality with formal languages
Josef Valvoda, Naomi Saphra, Jonathan Rawski, Adina Williams, and Ryan Cotterell · 2022
Cited alongside, same era.
The shape of learning curves: a review
Tom Viering and Marco Loog · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al · 2022
Cited alongside, same era.
What shapes the loss landscape of self-supervised learning?
Liu Ziyin, Ekdeep Singh Lubana, Masahito Ueda, and Hidenori Tanaka · 2022
Cited alongside, same era.
Grokking phase transitions in learning local rules with gradient descent, October 2022
Bojan Žunkovič and Enej Ilievski · 2022
Cited alongside, same era.
Later among the works it cites.
Grokking of Hierarchical Structure in Vanilla Transformers, May 2023
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D. Manning · 2023
Later among the works it cites.
Emergent Linear Representations in World Models of Self-Supervised Sequence Models, sep 2023
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task
Maya Okawa, Ekdeep Singh Lubana, Robert P Dick, and Hidenori Tanaka · 2023
Later among the works it cites.
Gpt-4, 2023
OpenAI · 2023
Later among the works it cites.
Blueprint for an AI Bill of Rights, 2023
The White House OSTP · 2023
Later among the works it cites.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Gautam Reddy · 2023
Later among the works it cites.
Are emergent abilities of Large Language Models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Later among the works it cites.
Emergent Deception and Emergent Optimization, feb 2023
Jacob Steinhardt · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, …, and Thomas Scialom · 2023
Later among the works it cites.
137 emergent abilities of large language models, 2022
Jason Wei · 2023
Later among the works it cites.
(Un)interpretability of Transformers: a case study with Dyck grammars
Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski · 2023
Later among the works it cites.
Skill-Mix: A flexible and expandable family of evaluations for AI models
Dingli Yu, Simran Kaur, Arushi Gupta, Jonah Brown-Cohen, Anirudh Goyal, and Sanjeev Arora · 2023
Later among the works it cites.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al · 2024
Closest in time.
Towards a theory of how the structure of language is acquired by deep neural networks
Francesco Cagnetta and Matthieu Wyart · 2024
Closest in time.
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt, and Naomi Saphra · 2024
Closest in time.
Proposal for a Regulation of the European Parliament and of the Council on Artificial Intelligence (Artificial Intelligence Act), 2024
Council of the European Union · 2024
Closest in time.
Hugo Cui, Freya Behrens, Florent Krzakala, and Lenka Zdeborová · 2024
Closest in time.
Understanding emergent abilities of language models from the loss perspective
Zhengxiao Du, Aohan Zeng, Yuxiao Dong, and Jie Tang · 2024
Closest in time.
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
Benjamin L Edelman, Ezra Edelman, Surbhi Goel, Eran Malach, and Nikolaos Tsilivis · 2024
Closest in time.
Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks
Tianyu He, Darshil Doshi, Aritra Das, and Andrey Gromov · 2024
Closest in time.
An exactly solvable model for emergence and scaling laws
Yoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee, and Ard Louis · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2024
Closest in time.