Fetching the paper…
Reading the bibliography…
Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of the model capabilities.
Contribution to the theory of ferromagnetism
Ernst Ising · 1925
Earlier work this paper cites.
Crystal statistics. i. a two-dimensional model with an order-disorder transition
Lars Onsager · 1944
Earlier work this paper cites.
Generalized linear models
John Ashworth Nelder and Robert WM Wedderburn · 1972
Earlier work this paper cites.
Toward a mean field theory for spin glasses
Giorgio Parisi · 1979
Earlier work this paper cites.
Order parameter for spin-glasses
Giorgio Parisi · 1983
Earlier work this paper cites.
Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications
Marc Mézard, Giorgio Parisi, and Miguel Angel Virasoro · 1987
Earlier work this paper cites.
On milman’s inequality and random subspaces which escape through a mesh in rn
Yehoram Gordon · 1988
Earlier work this paper cites.
Learning from examples in large neural networks
Haim Sompolinsky, Naftali Tishby, and H Sebastian Seung · 1990
Earlier work this paper cites.
First-order transition to perfect generalization in a neural network with binary synapses
Géza Györgyi · 1990
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
Learning a rule in a multilayer neural network
Henry Schwarze · 1993
Earlier work this paper cites.
The lasso risk for gaussian matrices
Mohsen Bayati and Andrea Montanari · 2011
Earlier work this paper cites.
The dynamics of message passing on dense graphs, with applications to compressed sensing
Mohsen Bayati and Andrea Montanari · 2011
Earlier work this paper cites.
State evolution for general approximate message passing algorithms, with applications to spatial coupling
Adel Javanmard and Andrea Montanari · 2013
Earlier work this paper cites.
Meshes that trap random subspaces
Mihailo Stojnic · 2013
Earlier work this paper cites.
Upper-bounding l1-optimization weak thresholds
Mihailo Stojnic · 2013
Earlier work this paper cites.
An iterative construction of solutions of the tap equations for the sherrington–kirkpatrick model
Erwin Bolthausen · 2014
Earlier work this paper cites.
The gaussian min-max theorem in the presence of convexity
Christos Thrampoulidis, Samet Oymak, and Babak Hassibi · 2014
Earlier work this paper cites.
Fixed points of generalized approximate message passing with arbitrary matrices
Sundeep Rangan, Philip Schniter, Erwin Riegler, Alyson K Fletcher, and Volkan Cevher · 2016
Earlier work this paper cites.
High dimensional robust m-estimation: Asymptotic variance via approximate message passing
David Donoho and Andrea Montanari · 2016
Earlier work this paper cites.
Statistical physics of inference: Thresholds and algorithms
Lenka Zdeborová and Florent Krzakala · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Phase transitions in restricted boltzmann machines with generic priors
Adriano Barra, Giuseppe Genovese, Peter Sollich, and Daniele Tantari · 2017
Earlier work this paper cites.
The committee machine: Computational to statistical gaps in learning a two-layers neural network
Benjamin Aubin, Antoine Maillard, Florent Krzakala, Nicolas Macris, Lenka Zdeborová, et al · 2018
Earlier work this paper cites.
Self-Attention with Relative Position Representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Optimal errors and phase transitions in high-dimensional generalized linear models
Jean Barbier, Florent Krzakala, Nicolas Macris, Léo Miolane, and Lenka Zdeborová · 2019
Cited alongside, same era.
Generalized linear models
Peter McCullagh · 2019
Cited alongside, same era.
Finding the needle in the haystack with convolutions: on the benefits of architectural bias
Stéphane d'Ascoli, Levent Sagun, Giulio Biroli, and Joan Bruna · 2019
Cited alongside, same era.
Phase retrieval in high dimensions: Statistical and computational phase transitions
Antoine Maillard, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Cited alongside, same era.
The role of regularization in classification of high-dimensional noisy gaussian mixture
A theory for emergence of complex skills in language models
Sanjeev Arora and Anirudh Goyal · 2023
Later among the works it cites.
Optimal inference of a generalised Potts model by single-layer transformers with factored attention
Riccardo Rende, Federica Gerace, Alessandro Laio, and Sebastian Goldt · 2023
Later among the works it cites.
Max-margin token selection in attention mechanism
Davoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, and Samet Oymak · 2023
Later among the works it cites.
Transformers as support vector machines
Davoud Ataee Tarzanagh, Yingcong Li, Christos Thrampoulidis, and Samet Oymak · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Francesca Mignacco, Florent Krzakala, Yue Lu, Pierfrancesco Urbani, and Lenka Zdeborova · 2020
Cited alongside, same era.
Generalization error in high-dimensional perceptrons: Approaching bayes error with convex optimization
Benjamin Aubin, Florent Krzakala, Yue Lu, and Lenka Zdeborová · 2020
Cited alongside, same era.
Generalization error of generalized linear models in high dimensions
Melikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan, and Alyson Fletcher · 2020
Cited alongside, same era.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Cited alongside, same era.
Learning gaussian mixtures with generalized linear models: Precise asymptotics in high-dimensions
Bruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco, Florent Krzakala, and Lenka Zdeborová · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Sander, and Yuanzhi Li · 2022
Cited alongside, same era.
Hongkang Li, Meng Wang, Sijia Liu, and Pin-Yu Chen · 2023
Later among the works it cites.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon S Du · 2023
Later among the works it cites.
Tianyu Guo, Wei Hu, Song Mei, Huan Wang, Caiming Xiong, Silvio Savarese, and Yu Bai · 2023
Later among the works it cites.
Transformers as algorithms: Generalization and stability in in-context learning
Yingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett · 2023
Later among the works it cites.
A mathematical perspective on transformers
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan, Payel Das, and Siva Reddy · 2023
Later among the works it cites.
Randomized positional encodings boost length generalization of transformers
Anian Ruoss, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Róbert Csordás, Mehdi Bennani, Shane Legg, and Joel Veness · 2023
Later among the works it cites.
Learning curves for the multi-class teacher–student perceptron
Elisabetta Cornacchia, Francesca Mignacco, Rodrigo Veiga, Cédric Gerbelot, Bruno Loureiro, and Lenka Zdeborová · 2023
Later among the works it cites.
Theoretical characterization of uncertainty in high-dimensional linear classification
Lucas Clarté, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2024
Closest in time.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression
Allan Raventós, Mansheej Paul, Feng Chen, and Surya Ganguli · 2024
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Gautam Reddy · 2024
Closest in time.
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2024
Closest in time.
What can a single attention layer learn? a study through the random features lens
Hengyu Fu, Tianyu Guo, Yu Bai, and Song Mei · 2024
Closest in time.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2024
Closest in time.
Transformers learn through gradual rank increase
Emmanuel Abbe, Samy Bengio, Enric Boix-Adsera, Etai Littwin, and Joshua Susskind · 2024
Closest in time.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2024
Closest in time.
High-dimensional asymptotics of denoising autoencoders
Hugo Cui and Lenka Zdeborová · 2024
Closest in time.
High-dimensional learning of narrow neural networks
Hugo Cui · 2024
Closest in time.