Fetching the paper…
Reading the bibliography…
Our understanding of the generalization capabilities of neural networks (NNs) is still incomplete.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 1901
Earlier work this paper cites.
Photograph enhancement by adaptive digital unsharp masking
Alan Martin Gilkes · 1974
Earlier work this paper cites.
The need for biases in learning generalizations
Tom M Mitchell · 1980
Earlier work this paper cites.
There exists a neural network that does not make avoidable mistakes
Gallant · 1988
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Discovering neural nets with low kolmogorov complexity and high generalization capability
Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, Andrew Y Ng, et al · 2013
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Inductive bias of deep convolutional networks through pooling geometry
Nadav Cohen and Amnon Shashua · 2016
Earlier work this paper cites.
Deep neural networks with random gaussian weights: A universal classification strategy?
Raja Giryes, Guillermo Sapiro, and Alex M Bronstein · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, JP Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Failures of gradient-based deep learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Earlier work this paper cites.
Input–output maps are strongly biased towards simple outputs
Kamaludin Dingle, Chico Q Camargo, and Ard A Louis · 2018
Earlier work this paper cites.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2018
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Earlier work this paper cites.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2018
Earlier work this paper cites.
Theory of deep learning III: the non-overfitting puzzle
Tomaso Poggio, Kenji Kawaguchi, Qianli Liao, Brando Miranda, Lorenzo Rosasco, Xavier Boix, Jack Hidary, and Hrushikesh Mhaskar · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
On the learning dynamics of deep neural networks
Remi Tachet, Mohammad Pezeshki, Samira Shabanian, Aaron Courville, and Yoshua Bengio · 2018
Earlier work this paper cites.
Deep image prior
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2018
Cited alongside, same era.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Perez, Chico Q Camargo, and Ard A Louis · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
Random deep neural networks are biased towards simple functions
Giacomo De Palma, Bobak Kiani, and Seth Lloyd · 2019
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Cited alongside, same era.
Which shortcut cues will dnns choose? a study from the parameter-space perspective
Luca Scimeca, Seong Joon Oh, Sanghyuk Chun, Michael Poli, and Sangdoo Yun · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David GT Barrett, and Soham De · 2021
Later among the works it cites.
Damien Teney, Ehsan Abbasnejad, Simon Lucey, and Anton van den Hengel · 2021
Later among the works it cites.
Good classifiers are abundant in the interpolating regime
Ryan Theisen, Jason Klusowski, and Michael Mahoney · 2021
Later among the works it cites.
Simplicity bias in transformers and their ability to learn sparse boolean functions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chris Mingard, Joar Skalse, Guillermo Valle-Pérez, David Martínez-Rubio, Vladimir Mikulik, and Ard A Louis · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
The bitter lesson
Richard Sutton · 2019
Cited alongside, same era.
A fine-grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Cited alongside, same era.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Cited alongside, same era.
Satwik Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom · 2022
Later among the works it cites.
Loss landscapes are all you need: Neural network generalization can be explained without the implicit bias of gradient descent
Ping-yeh Chiang, Renkun Ni, David Yu Miller, Arpit Bansal, Jonas Geiping, Micah Goldblum, and Tom Goldstein · 2022
Later among the works it cites.
The spectral bias of polynomial neural networks
Moulik Choraria, Leello Tadesse Dadi, Grigorios Chrysos, Julien Mairal, and Volkan Cevher · 2022
Later among the works it cites.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, et al · 2022
Later among the works it cites.
Why neural networks find simple solutions: The many regularizers of geometric complexity
Benoit Dherin, Michael Munn, Mihaela Rosca, and David Barrett · 2022
Later among the works it cites.
Activation functions in deep learning: A comprehensive survey and benchmark
Shiv Ram Dubey, Satish Kumar Singh, and Bidyut Baran Chaudhuri · 2022
Later among the works it cites.
Implicit bias in leaky relu networks trained on high-dimensional data
Spencer Frei, Gal Vardi, Peter L Bartlett, Nathan Srebro, and Wei Hu · 2022
Later among the works it cites.
On the activation function dependence of the spectral bias of neural networks
Qingguo Hong, Jonathan W Siegel, Qinyang Tan, and Jinchao Xu · 2022
Later among the works it cites.
Beyond periodicity: Towards a unifying framework for activations in coordinate-MLPs
Sameera Ramasinghe and Simon Lucey · 2022
Later among the works it cites.
Reverse engineering the neural tangent kernel
James Benjamin Simon, Sajant Anand, and Mike Deweese · 2022
Later among the works it cites.
Property unlearning: A defense strategy against property inference attacks
Joshua Stock, Jens Wettlaufer, Daniel Demmler, and Hannes Federrath · 2022
Later among the works it cites.
Predicting is not understanding: Recognizing and addressing underspecification in machine learning
Damien Teney, Maxime Peyrard, and Ehsan Abbasnejad · 2022
Later among the works it cites.
Neural fields in visual computing and beyond
Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, and Srinath Sridhar · 2022
Later among the works it cites.
Generalization on the unseen, logic reasoning and degree curriculum
Emmanuel Abbe, Samy Bengio, Aryo Lotfi, and Kevin Rizk · 2023
Later among the works it cites.
Scaling MLPs: A tale of inductive bias
Gregor Bachmann, Sotiris Anagnostidis, and Thomas Hofmann · 2023
Later among the works it cites.
Model-agnostic measure of generalization difficulty
Akhilan Boopathy, Kevin Liu, Jaedong Hwang, Shu Ge, Asaad Mohammedsaleh, and Ila R Fiete · 2023
Later among the works it cites.
Initial guessing bias: How untrained networks favor some classes
Emanuele Francazi, Aurelien Lucchi, and Marco Baity-Jesi · 2023
Later among the works it cites.
Micah Goldblum, Marc Finzi, Keefer Rowan, and Andrew Gordon Wilson · 2023
Later among the works it cites.
On the foundations of shortcut learning
Katherine L Hermann, Hossein Mobahi, Thomas Fel, and Michael C Mozer · 2023
Later among the works it cites.
Special properties of gradient descent with large learning rates
Amirkeivan Mohtashami, Martin Jaggi, and Sebastian U Stich · 2023
Later among the works it cites.
Wire: Wavelet implicit neural representations
Vishwanath Saragadam, Daniel LeJeune, Jasper Tan, Guha Balakrishnan, Ashok Veeraraghavan, and Richard G Baraniuk · 2023
Later among the works it cites.
On the implicit bias in deep-learning algorithms
Gal Vardi · 2023
Later among the works it cites.
Do we always need the simplicity bias? Looking for optimal inductive biases in the wild
Damien Teney, Liangze Jiang, Florin Gogianu, and Ehsan Abbasnejad · 2025
Closest in time.