Fetching the paper…
Reading the bibliography…
The remarkable capability of over-parameterised neural networks to generalise effectively has been explained by invoking a ``simplicity bias'': neural networks prevent overfitting by initially learning simple classifiers before progressing to more complex, non-linear functions.
Compatible conditional distributions
Barry C. Arnold and S. James Press · 1989
Earlier work this paper cites.
Exact Solution for On-Line Learning in Multilayer Neural Networks
D. Saad and S.A. Solla · 1995
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A spectral approach to generalization and optimization in neural networks
F. Farnia, J. Zhang, and D. Tse · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P. Nguyen · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2018
Earlier work this paper cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle Pérez, Chico Q. Camargo, and Ard A. Louis · 2019
Earlier work this paper cites.
SGD on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin L. Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Earlier work this paper cites.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A. Hamprecht, Yoshua Bengio, and Aaron C. Courville · 2019
Earlier work this paper cites.
A mathematical theory of semantic development in deep neural networks
A. Saxe, J.L. McClelland, and S. Ganguli · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
M.S. Advani, A.M. Saxe, and H. Sompolinsky · 2020
Earlier work this paper cites.
Interpreting potts and transformer protein models through the lens of simplified attention
Nicholas Bhattacharya, Neil Thomas, Roshan Rao, Justas Dauparas, Peter K. Koo, David Baker, Yun S. Song, and Sergey Ovchinnikov · 2020
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing, 2020
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert, 2020
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture, 2020
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu · 2020
Cited alongside, same era.
Probing across time: What does roberta know and when?
Zeyu Liu, Yizhong Wang, Jungo Kasai, Hannaneh Hajishirzi, and Noah A Smith · 2021
Cited alongside, same era.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Cited alongside, same era.
How to train bert with an academic budget, 2021
Peter Izsak, Moshe Berchansky, and Omer Levy · 2021
Cited alongside, same era.
Implicit bias of MSE gradient optimization in underparameterized neural networks
Transformer variational wave functions for frustrated quantum spin systems
Luciano Loris Viteritti, Riccardo Rende, and Federico Becca · 2023
Later among the works it cites.
Transformer wave function for the shastry-sutherland model: emergence of a spin-liquid phase
Luciano Loris Viteritti, Riccardo Rende, Alberto Parola, Sebastian Goldt, and Federico Becca · 2023
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbé, Enric Boix Adserà, and Theodor Misiakiewicz · 2023
Later among the works it cites.
Learning two-layer neural networks, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benjamin Bowman and Guido Montúfar · 2022
Cited alongside, same era.
Spectral bias outside the training set for deep networks in the kernel regime
Benjamin Bowman and Guido Montúfar · 2022
Cited alongside, same era.
Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs
Etienne Boursier, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Cited alongside, same era.
Exposing the implicit energy networks behind masked language models via metropolis–hastings, 2022
Kartik Goyal, Chris Dyer, and Taylor Berg-Kirkpatrick · 2022
Cited alongside, same era.
Simplicity bias in transformers and their ability to learn sparse boolean functions
Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom · 2022
Cited alongside, same era.
The grammar-learning trajectories of neural language models
Leshem Choshen, Guy Hacohen, Daphna Weinshall, and Omri Abend · 2022
Cited alongside, same era.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Cited alongside, same era.
Raphaël Berthier, Andrea Montanari, and Kangjie Zhou · 2023
Later among the works it cites.
Sliding down the stairs: How correlated latent variables accelerate learning with neural networks
Lorenzo Bardone and Sebastian Goldt · 2024
Closest in time.
Neural networks learn statistics of increasing complexity
Nora Belrose, Quintin Pope, Lucia Quirke, Alex Mallen, and Xiaoli Fern · 2024
Closest in time.
Mapping of attention mechanisms to a generalized potts model
Riccardo Rende, Federica Gerace, Alessandro Laio, and Sebastian Goldt · 2024
Closest in time.
Transformers learn through gradual rank increase
Emmanuel Abbe, Samy Bengio, Enric Boix-Adsera, Etai Littwin, and Joshua Susskind · 2024
Closest in time.
Simplicity bias of transformers to learn low sensitivity functions
Bhavya Vasudeva, Deqing Fu, Tianyi Zhou, Elliott Kau, Youqi Huang, and Vatsal Sharan · 2024
Closest in time.
How deep neural networks learn compositional data: The random hierarchy model
Francesco Cagnetta, Leonardo Petrini, Umberto M Tomasini, Alessandro Favero, and Matthieu Wyart · 2024
Closest in time.
Towards a theory of how the structure of language is acquired by deep neural networks
Francesco Cagnetta and Matthieu Wyart · 2024
Closest in time.
How transformers learn structured data: insights from hierarchical filtering
Jérôme Garnier-Brun, Marc Mézard, Emanuele Moscato, and Luca Saglietti · 2024
Closest in time.
A simple linear algebra identity to optimize large-scale neural network quantum states
Riccardo Rende, Luciano Loris Viteritti, Lorenzo Bardone, Federico Becca, and Sebastian Goldt · 2024
Closest in time.
Fine-tuning neural network quantum states
Riccardo Rende, Sebastian Goldt, Federico Becca, and Luciano Loris Viteritti · 2024
Closest in time.
Smoothing the landscape boosts the signal for sgd: Optimal sample complexity for learning single index models
Alex Damian, Eshaan Nichani, Rong Ge, and Jason D Lee · 2024
Closest in time.