Fetching the paper…
Reading the bibliography…
The power of neural networks lies in their ability to generalize to unseen data, yet the underlying reasons for this phenomenon remain elusive.
On Milman’s inequality and random subspaces which escape through a mesh in ℝ \mathbb{R} n
Yehoram Gordon · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
G. Cybenko · 1989
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Almost linear vc dimension bounds for piecewise polynomial networks
Peter L. Bartlett, Vitaly Maiorov, and Ron Meir · 1998
Earlier work this paper cites.
Some pac-bayesian theorems
David A. McAllester · 1998
Earlier work this paper cites.
Pac-bayesian model averaging
David A. McAllester · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2003
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey E. Hinton · 2008
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y. Ng · 2011
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Pengtao Xie, Yuntian Deng, and Eric Xing · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
On the depth of deep neural networks: A theoretical view
Shizhao Sun, Wei Chen, Liwei Wang, Xiaoguang Liu, and Tie-Yan Liu · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2016
Cited alongside, same era.
Certified defenses for data poisoning attacks
Jacob Steinhardt, Pang Wei W Koh, and Percy S Liang · 2017
Cited alongside, same era.
Nearly-tight VC-dimension bounds for piecewise linear neural networks
Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Later among the works it cites.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James Brecht · 2018
Later among the works it cites.
On the importance of single directions for generalization
Ari S. Morcos, David G.T. Barrett, Neil C. Rabinowitz, and Matthew Botvinick · 2018
Later among the works it cites.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2018
Later among the works it cites.
Data-dependent stability of stochastic gradient descent
Ilja Kuzborskij and Christoph H. Lampert · 2018
Later among the works it cites.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast rates for empirical risk minimization of strict saddle problems
Alon Gonen and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Generalization error of invariant classifiers
Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel Rodrigues · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
P Chaudhari, Anna Choromanska, S Soatto, Yann LeCun, C Baldassi, C Borgs, J Chayes, Levent Sagun, and R Zecchina · 2017
Cited alongside, same era.
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Later among the works it cites.
Identifying generalization properties in neural networks
Huan Wang, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Later among the works it cites.
Transferable clean-label poisoning attacks on deep neural nets
Chen Zhu, W Ronny Huang, Hengduo Li, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2019
Closest in time.