Fetching the paper…
Reading the bibliography…
A key property of deep neural networks (DNNs) is their ability to learn new features during training.
Optimal storage properties of neural network models
E. Gardner and B. Derrida · 1988
Earlier work this paper cites.
First-order transition to perfect generalization in a neural network with binary synapses
Géza Györgyi · 1990
Earlier work this paper cites.
Statistical mechanics of learning from examples
H. S. Seung, H. Sompolinsky, and N. Tishby · 1992
Earlier work this paper cites.
Exact solution for on-line learning in multilayer neural networks
David Saad and Sara A. Solla · 1995
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
Computing with infinite networks
Christopher Williams · 1996
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Carl Edward Rasmussen and Christopher K. I. Williams · 2005
Earlier work this paper cites.
MCMC using hamiltonian dynamics
Radford M Neal et al · 2011
Earlier work this paper cites.
Statistical physics
David Tong · 2011
Earlier work this paper cites.
Statistical Physics: Volume 5 , volume 5
Lev Davidovich Landau and Evgenii Mikhailovich Lifshitz · 2013
Earlier work this paper cites.
Exploring one pass learning for deep neural network training with averaged stochastic gradient descent
Zhao You, Xiaorui Wang, and Bo Xu · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D. Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Deep learning and the information bottleneck principle, 2015
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
A survey of transfer learning
Karl Weiss, Taghi M Khoshgoftaar, and DingDing Wang · 2016
Cited alongside, same era.
Interpretability of deep learning models: A survey of results
Supriyo Chakraborty, Richard Tomsett, Ramya Raghavendra, Daniel Harborne, Moustafa Alzantot, Federico Cerutti, Mani Srivastava, Alun Preece, Simon Julier, Raghuveer M. Rao, Troy D. Kelley, Dave Braines, Murat Sensoy, Christopher J. Willis, and Prudhvi Gurram · 2017
Cited alongside, same era.
Nonasymptotic convergence analysis for the unadjusted langevin algorithm
Alain Durmus and Eric Moulines · 2017
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Predicting the outputs of finite deep neural networks trained with noisy gradients
Gadi Naveh, Oded Ben David, Haim Sompolinsky, and Zohar Ringel · 2021
Later among the works it cites.
Statistical mechanics of deep learning beyond the infinite-width limit
S Ariosto, R Pacelli, M Pastore, F Ginelli, M Gherardi, and P Rotondo · 2022
Later among the works it cites.
Learning single-index models with shallow neural networks
Alberto Bietti, Joan Bruna, Clayton Sanford, and Min Jae Song · 2022
Later among the works it cites.
Self-consistent dynamical field theory of kernel evolution in wide neural networks
Blake Bordelon and Cengiz Pehlevan · 2022
Later among the works it cites.
Neural networks efficiently learn low-dimensional representations with sgd
Alireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas, and Murat A Erdogdu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Gamarnik, Eren C Kızıldağ, and Ilias Zadik · 2019
Cited alongside, same era.
Yan Shuo Tan and Roman Vershynin · 2019
Cited alongside, same era.
Optimization and generalization of shallow neural networks with quadratic activation functions
Stefano Sarao Mannelli, Eric Vanden-Eijnden, and Lenka Zdeborová · 2020
Cited alongside, same era.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Cited alongside, same era.
Learning curves for overparametrized deep neural networks: A field theory perspective
Omry Cohen, Or Malka, and Zohar Ringel · 2021
Cited alongside, same era.
Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization
Qianyi Li and Haim Sompolinsky · 2021
Cited alongside, same era.
A self consistent theory of gaussian processes captures feature learning effects in finite cnns
Gadi Naveh and Zohar Ringel · 2021
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Grokking phase transitions in learning local rules with gradient descent
Bojan Žunkovič and Enej Ilievski · 2022
Later among the works it cites.
Luca Arnaboldi, Ludovic Stephan, Florent Krzakala, and Bruno Loureiro · 2023
Closest in time.
Andrey Gromov · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability, 2023
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Separation of scales and a thermodynamic description of feature learning in some CNNs
Inbar Seroussi, Gadi Naveh, and Zohar Ringel · 2023
Closest in time.
Explaining grokking through circuit efficiency
Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar · 2023
Closest in time.