Fetching the paper…
Reading the bibliography…
We approach the problem of implicit regularization in deep learning from a geometrical viewpoint.
Pathological spectra of the fisher information metric and its variants in deep neural networks
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 1910
Earlier work this paper cites.
Neural networks and the bias/variance dilemma
Stuart Geman, Elie Bienenstock, and René Doursat · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Kernel-dependent support vector error bounds
B. Schölkopf, J. Shawe-Taylor, AJ. Smola, and RC. Williamson · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
On kernel-target alignment
Nello Cristianini, John Shawe-Taylor, André Elisseeff, and Jaz S. Kandola · 2002
Earlier work this paper cites.
Learning eigenfunctions links spectral embedding and kernel PCA
Yoshua Bengio, Olivier Delalleau, Nicolas Le Roux, Jean-François Paiement, Pascal Vincent, and Marie Ouimet · 2004
Earlier work this paper cites.
Spectral properties of the kernel matrix and their relation to kernel methods in machine learning
Mikio L Braun · 2005
Earlier work this paper cites.
Measuring statistical dependence with hilbert-schmidt norms, 2005
Arthur Gretton, Olivier Bousquet, Alexander Smola, and Bernhard Schölkopf · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Olivier Roy and Martin Vetterli · 2007
Earlier work this paper cites.
The elements of statistical learning: data mining, inference and prediction
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
On the universality of online mirror descent
Nati Srebro, Karthik Sridharan, and Ambuj Tewari · 2011
Earlier work this paper cites.
Algorithms for learning kernels based on centered alignment
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh · 2012
Earlier work this paper cites.
Fundamentals of Differential Geometry
S. Lang · 2012
Earlier work this paper cites.
Foundations of Machine Learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2012
Earlier work this paper cites.
Probability in Banach Spaces: Isoperimetry and Processes
M. Ledoux and M. Talagrand · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural network
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Information Geometry and Its Applications , volume 194
Shun-Ichi Amari · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Cited alongside, same era.
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stephano Spigler, Arthur Jacot, and Matthieu Wyart · 2019
Later among the works it cites.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Later among the works it cites.
Universal statistics of fisher information in deep neural networks: Mean field approach
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2019
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2019
Later among the works it cites.
Harmless interpolation of noisy data in regression
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spectrally-normalized margin bounds for neural networks
Peter L. Bartlett, Dylan J. Foster, and Matus Telgarsky · 2017
Cited alongside, same era.
Train longer, generalize better: Closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
Deep learning for classical japanese literature
Tarin Clanuwat, Mikel Bober-Irizar, Asanobu Kitamoto, Alex Lamb, Kazuaki Yamamoto, and David Ha · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Later among the works it cites.
Training behavior of deep neural network in frequency domain
Zhi-Qin John Xu, Yaoyu Zhang, and Yanyang Xiao · 2019
Later among the works it cites.
A fine grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
Later among the works it cites.
Coherent gradients: An approach to understanding generalization in gradient descent-based optimization
Satrajit Chatterjee · 2020
Closest in time.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Closest in time.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2020
Closest in time.
Neural spectrum alignment: Empirical study
D. Kopitkov and V. Indelman · 2020
Closest in time.
Classification vs regression in overparameterized regimes: Does the loss function matter?
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Sahai, Hsu, and Anant Sahai · 2020
Closest in time.
Geometric compression of invariant manifolds in neural nets
Jonas Paccolat, Leonardo Petrini, Mario Geiger, Kevin Tyloo, and Matthieu Wyart · 2020
Closest in time.
An investigation of why overparameterization exacerbates spurious correlations
Shiori Sagawa, Aditi Raghunathan, Pang Wei Koh, and Percy Liang · 2020
Closest in time.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Closest in time.
Weighted optimization: better generalization by smoother interpolation
Yuege Xie, Rachel Ward, Holger Rauhut, and Chou Hung-Hsu · 2020
Closest in time.
NNGeometry: Easy and Fast Fisher Information Matrices and Neural Tangent Kernels in PyTorch, February 2021
Thomas George · 2021
Closest in time.