Fetching the paper…
Reading the bibliography…
There is mounting evidence of emergent phenomena in the capabilities of deep learning methods as we scale up datasets, model sizes, and training times.
The accuracy of the gaussian approximation to the sum of independent variates
Andrew C Berry · 1941
Earlier work this paper cites.
On the Liapunov limit error in the theory of probability
Carl-Gustav Esseen · 1942
Earlier work this paper cites.
Inequalities
Godfrey Harold Hardy, John Edensor Littlewood, George Pólya, György Pólya, et al · 1952
Earlier work this paper cites.
Correlation properties of cyclic sequences
Robert C Titsworth · 1962
Earlier work this paper cites.
Three unfinished works on the optimal storage capacity of networks
Elizabeth Gardner and Bernard Derrida · 1989
Earlier work this paper cites.
A hard-core predicate for all one-way functions
Oded Goldreich and Leonid A Levin · 1989
Earlier work this paper cites.
Bounds on the learning capacity of some multi-layer networks
GJ Mitchison and RM Durbin · 1989
Earlier work this paper cites.
Memorization without generalization in a multilayered neural network
D Hansel, G Mato, and C Meunier · 1992
Earlier work this paper cites.
Learning decision trees using the fourier spectrum
Eyal Kushilevitz and Yishay Mansour · 1993
Earlier work this paper cites.
The statistical mechanics of learning a rule
Timothy LH Watkin, Albrecht Rau, and Michael Biehl · 1993
Earlier work this paper cites.
Perfect loss of generalization due to noise in k= 2 parity machines
Y Kabashima · 1994
Earlier work this paper cites.
Learning and generalization in a two-layer neural network: The role of the vapnik-chervonvenkis dimension
Manfred Opper · 1994
Earlier work this paper cites.
On-line learning in parity machines
Roberta Simonetti and Nestor Caticha · 1996
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Mutual learning in a tree parity machine and its application to cryptography
Michal Rosen-Zvi, Einat Klein, Ido Kanter, and Wolfgang Kinzel · 2002
Earlier work this paper cites.
More on average case vs approximation complexity
Michael Alekhnovich · 2003
Earlier work this paper cites.
Learning intersections and thresholds of halfspaces
Adam R Klivans, Ryan O’Donnell, and Rocco A Servedio · 2004
Earlier work this paper cites.
Cryptography with constant computational overhead
Yuval Ishai, Eyal Kushilevitz, Rafail Ostrovsky, and Amit Sahai · 2008
Earlier work this paper cites.
Fast cryptographic primitives and circular-secure encryption based on hard learning problems
Benny Applebaum, David Cash, Chris Peikert, and Amit Sahai · 2009
Earlier work this paper cites.
Public-key cryptography from different assumptions
Benny Applebaum, Boaz Barak, and Avi Wigderson · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Embedding hard learning problems into gaussian space
Adam Klivans and Pravesh Kothari · 2014
Earlier work this paper cites.
Analysis of Boolean functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Time-space hardness of learning sparse parities
Gillat Kol, Ran Raz, and Avishay Tal · 2017
Cited alongside, same era.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Approximation schemes for relu regression
Ilias Diakonikolas, Surbhi Goel, Sushrut Karmalkar, Adam R Klivans, and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Agnostic learning of a single neuron with gradient descent
Spencer Frei, Yuan Cao, and Quanquan Gu · 2020
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Later among the works it cites.
Approximate is good enough: Probabilistic variants of dimensional and margin complexity
Pritish Kamath, Omar Montasser, and Nathan Srebro · 2020
Later among the works it cites.
Learning a single neuron with gradient methods
Gilad Yehudai and Shamir Ohad · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
What can ResNet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Greedy layerwise learning can scale to imagenet
Eugene Belilovsky, Michael Eickenberg, and Edouard Oyallon · 2019
Cited alongside, same era.
Xor codes and sparse learning parity with noise
Andrej Bogdanov, Manuel Sabin, and Prashant Nalini Vasudevan · 2019
Cited alongside, same era.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2019
Cited alongside, same era.
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
On the power of differentiable learning versus PAC and SQ learning
Emmanuel Abbe, Pritish Kamath, Eran Malach, Colin Sandon, and Nathan Srebro · 2021
Later among the works it cites.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2021
Later among the works it cites.
Pretrained transformers as universal computation engines
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch · 2021
Later among the works it cites.
Quantifying the benefit of using differentiable learning over tangent kernels
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro · 2021
Later among the works it cites.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Later among the works it cites.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Zhenmei Shi, Junyi Wei, and Yingyu Liang · 2021
Later among the works it cites.
An initial alignment between neural network and target is needed for gradient descent to learn
Emmanuel Abbe, Elisabetta Cornacchia, Jan Hązła, and Christopher Marquis · 2022
Closest in time.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur · 2022
Closest in time.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
Neural networks can learn representations with gradient descent
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Closest in time.
Random feature amplification: Feature learning and generalization in neural networks
Spencer Frei, Niladri S Chatterji, and Peter L Bartlett · 2022
Closest in time.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Closest in time.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Closest in time.
When hardness of approximation meets hardness of learning
Eran Malach and Shai Shalev-Shwartz · 2022
Closest in time.
A mechanistic interpretability analysis of grokking
Neel Nanda and Tom Lieberum · 2022
Closest in time.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Closest in time.
Feature selection with gradient descent on two-layer networks in low-rotation regimes
Matus Telgarsky · 2022
Closest in time.