Fetching the paper…
Reading the bibliography…
Overparameterised deep neural networks (DNNs) are highly expressive and so can, in principle, generate almost any function that fits a training dataset with zero error.
Greg Yang · 1902
Earlier work this paper cites.
A formal theory of inductive inference. part i
Ray J Solomonoff · 1964
Earlier work this paper cites.
Laws of information conservation (nongrowth) and aspects of the foundation of probability theory
L.A. Levin · 1974
Earlier work this paper cites.
Occam’s razor
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth · 1987
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Classification of radar returns from the ionosphere using neural networks
Vincent G Sigillito, Simon P Wing, Larrie V Hutton, and Kile B Baker · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Bayesian interpolation
David JC MacKay · 1992
Earlier work this paper cites.
Priors for infinite networks (tech. rep. no. crg-tr-94-1)
Radford M Neal · 1994
Earlier work this paper cites.
The relationship between PAC, the statistical physics framework, the bayesian framework, and the vc framework
David H Wolpert and R Waters · 1994
Earlier work this paper cites.
Introduction to gaussian processes
David JC Mackay · 1998
Earlier work this paper cites.
The role of occam’s razor in knowledge discovery
Pedro Domingos · 1999
Earlier work this paper cites.
Object recognition with gradient-based learning
Yann LeCun, Patrick Haffner, Léon Bottou, and Yoshua Bengio · 1999
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
High-dimensional data analysis: The curses and blessings of dimensionality
David L Donoho et al · 2000
Earlier work this paper cites.
Occam’s razor
Carl Edward Rasmussen and Zoubin Ghahramani · 2001
Earlier work this paper cites.
Information theory, inference and learning algorithms
David JC MacKay · 2003
Earlier work this paper cites.
Energy landscapes: Applications to clusters, biomolecules and glasses
David Wales et al · 2003
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2004
Earlier work this paper cites.
Intrinsic dimensionality estimation of submanifolds in rd
Matthias Hein and Jean-Yves Audibert · 2005
Earlier work this paper cites.
Power-law distributions for the areas of the basins of attraction on a potential energy landscape
Claire P Massen and Jonathan PK Doye · 2007
Earlier work this paper cites.
An introduction to Kolmogorov complexity and its applications
M. Li and P.M.B. Vitanyi · 2008
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
A philosophical treatise of universal induction
Samuel Rathmanner and Marcus Hutter · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Radford M Neal · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Adam: Amethod for stochastic optimization
Diederik P Kingma and Jimmy Lei Ba · 2014
Earlier work this paper cites.
The arrival of the frequent: how bias in genotype-phenotype maps can steer populations to local optima
Steffen Schaper and Ard A Louis · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
The structure of the genotype–phenotype map strongly constrains the evolution of non-coding rna
Kamaludin Dingle, Steffen Schaper, and Ard A Louis · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Ockham’s razors
Elliott Sober · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
Solomonoff prediction and occam’s razor
Tom F Sterkenburg · 2016
Deep neural networks as gaussian processes
Jaehoon Lee, Jascha Sohl-dickstein, Jeffrey Pennington, Roman Novak, Sam Schoenholz, and Yasaman Bahri · 2018
Later among the works it cites.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Later among the works it cites.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2018
Later among the works it cites.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A Hamprecht, Yoshua Bengio, and Aaron Courville · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Cited alongside, same era.
Energy landscapes for machine learning
Andrew J Ballard, Ritankar Das, Stefano Martiniani, Dhagash Mehta, Levent Sagun, Jacob D Stevenson, and David J Wales · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Nonlinear dynamics and chaos with student solutions manual: With applications to physics, biology, chemistry, and engineering
Steven H Strogatz · 2018
Later among the works it cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Pérez, Chico Q Camargo, and Ard A Louis · 2018
Later among the works it cites.
Energy–entropy competition and the effectiveness of stochastic gradient descent in machine learning
Yao Zhang, Andrew M Saxe, Madhu S Advani, and Alpha A Lee · 2018
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Later among the works it cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Later among the works it cites.
Comparing dynamics: Deep neural networks versus glassy systems
Marco Baity-Jesi, Levent Sagun, Mario Geiger, Stefano Spigler, Gérard Ben Arous, Chiara Cammarota, Yann LeCun, Matthieu Wyart, and Giulio Biroli · 2019
Later among the works it cites.
On empirical comparisons of optimizers for deep learning
Dami Choi, Christopher J Shallue, Zachary Nado, Jaehoon Lee, Chris J Maddison, and George E Dahl · 2019
Later among the works it cites.
Learning curves for deep neural networks: a gaussian field theory perspective
Omry Cohen, Or Malka, and Zohar Ringel · 2019
Later among the works it cites.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2019
Later among the works it cites.
Modelling the influence of data structure on learning in neural networks
Sebastian Goldt, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2019
Later among the works it cites.
Universal function approximation by deep neural nets with bounded width and relu activations
Boris Hanin · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
Neural networks are a priori biased towards boolean functions with low entropy
Chris Mingard, Joar Skalse, Guillermo Valle-Pérez, David Martínez-Rubio, Vladimir Mikulik, and Ard A Louis · 2019
Later among the works it cites.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Later among the works it cites.
On the information bottleneck theory of deep learning
Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox · 2019
Later among the works it cites.
Asymptotic learning curves of kernel methods: empirical data vs teacher-student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2019
Later among the works it cites.
How noise affects the hessian spectrum in overparameterized neural networks
Mingwei Wei and David J Schwab · 2019
Later among the works it cites.
A fine-grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a PAC-bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P. Adams, and Peter Orbanz · 2019
Later among the works it cites.
Geometry of energy landscapes and the optimizability of deep neural networks
Simon Becker, Yao Zhang, et al · 2020
Closest in time.
Can implicit bias explain generalization? stochastic convex optimization as a case study
Assaf Dauber, Meir Feder, Tomer Koren, and Roi Livni · 2020
Closest in time.
Generic predictions of output probability based on complexities of inputs and outputs
Kamaludin Dingle, Guillermo Valle Pérez, and Ard A Louis · 2020
Closest in time.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Closest in time.
Predicting the outputs of finite networks trained with noisy gradients
Gadi Naveh, Oded Ben-David, Haim Sompolinsky, and Zohar Ringel · 2020
Closest in time.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
Closest in time.
Complexity control by gradient descent in deep networks
Tomaso Poggio, Qianli Liao, and Andrzej Banburski · 2020
Closest in time.
How good is the bayes posterior in deep neural networks really?
Florian Wenzel, Kevin Roth, Bastiaan S Veeling, Jakub Świątkowski, Linh Tran, Stephan Mandt, Jasper Snoek, Tim Salimans, Rodolphe Jenatton, and Sebastian Nowozin · 2020
Closest in time.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew Gordon Wilson and Pavel Izmailov · 2020
Closest in time.