Fetching the paper…
Reading the bibliography…
Mathematically characterizing the implicit regularization induced by gradient-based optimization is a longstanding pursuit in the theory of deep learning.
Free states of the canonical anticommutation relations
Robert T Powers and Erling Størmer · 1970
Earlier work this paper cites.
Tensor rank is np-complete
Johan Håstad · 1990
Earlier work this paper cites.
A primer of real analytic functions
Steven G Krantz and Harold R Parks · 2002
Earlier work this paper cites.
A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization
Samuel Burer and Renato DC Monteiro · 2003
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Olivier Roy and Martin Vetterli · 2007
Earlier work this paper cites.
Lectures on analytic differential equations , volume 86
Yulij Ilyashenko and Sergei Yakovenko · 2008
Earlier work this paper cites.
Perturbation bounds for determinants and characteristic polynomials
Ilse CF Ipsen and Rizwana Rehman · 2008
Earlier work this paper cites.
Exact matrix completion via convex optimization
Emmanuel J Candès and Benjamin Recht · 2009
Earlier work this paper cites.
Tensor decompositions and applications
Tamara G Kolda and Brett W Bader · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Scalable tensor factorizations for incomplete data
Evrim Acar, Daniel M Dunlavy, Tamara G Kolda, and Morten Mørup · 2011
Earlier work this paper cites.
Null space conditions and thresholds for rank minimization
Benjamin Recht, Weiyu Xu, and Babak Hassibi · 2011
Earlier work this paper cites.
Matrix computations , volume 3
Gene H Golub and Charles F Van Loan · 2012
Earlier work this paper cites.
Tensor spaces and numerical tensor calculus , volume 42
Wolfgang Hackbusch · 2012
Earlier work this paper cites.
Tensor factorization using auxiliary information
Atsuhiro Narita, Kohei Hayashi, Ryota Tomioka, and Hisashi Kashima · 2012
Earlier work this paper cites.
Ordinary differential equations and dynamical systems , volume 140
Gerald Teschl · 2012
Earlier work this paper cites.
Perturbation theory for linear operators , volume 132
Tosio Kato · 2013
Earlier work this paper cites.
Tensor decompositions for learning latent variable models
Animashree Anandkumar, Rong Ge, Daniel Hsu, Sham M Kakade, and Matus Telgarsky · 2014
Earlier work this paper cites.
Simnets: A generalization of convolutional networks
Nadav Cohen and Amnon Shashua · 2014
Earlier work this paper cites.
Provable tensor factorization with missing data
Prateek Jain and Sewoong Oh · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Convolutional rectifier networks as generalized tensor decompositions
Nadav Cohen and Amnon Shashua · 2016
Earlier work this paper cites.
An overview of low-rank matrix recovery from incomplete observations
Mark A Davenport and Justin Romberg · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Parallel algorithms for tensor completion in the cp format
Lars Karlsson, Daniel Kressner, and André Uschmajew · 2016
Earlier work this paper cites.
Tensorial mixture models
Or Sharir, Ronen Tamari, Nadav Cohen, and Amnon Shashua · 2016
Cited alongside, same era.
Low-rank solutions of linear matrix equations via procrustes flow
Stephen Tu, Ross Boczar, Max Simchowitz, Mahdi Soltanolkotabi, and Ben Recht · 2016
Cited alongside, same era.
Smooth parafac decomposition for tensor completion
Tatsuya Yokota, Qibin Zhao, and Andrzej Cichocki · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Cited alongside, same era.
Inductive bias of deep convolutional networks through pooling geometry
Nadav Cohen and Amnon Shashua · 2017
Cited alongside, same era.
Analysis and design of convolutional networks via hierarchical tensor decompositions
Nonconvex low-rank tensor completion from noisy data
Changxiao Cai, Gen Li, H Vincent Poor, and Yuxin Chen · 2019
Later among the works it cites.
Nonconvex optimization meets low-rank matrix factorization: An overview
Yuejie Chi, Yue M Lu, and Yuxin Chen · 2019
Later among the works it cites.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Later among the works it cites.
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup
Sebastian Goldt, Madhu Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová · 2019
Later among the works it cites.
Sgd on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nadav Cohen, Or Sharir, Yoav Levine, Ronen Tamari, David Yakira, and Amnon Shashua · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Implicit regularization in deep learning
Behnam Neyshabur · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
On polynomial time methods for exact low rank tensor completion
Dong Xia and Ming Yuan · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Andrew K Lampinen and Surya Ganguli · 2019
Later among the works it cites.
Quantum entanglement in deep learning architectures
Yoav Levine, Or Sharir, Nadav Cohen, and Amnon Shashua · 2019
Later among the works it cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
On the spectral bias of deep neural networks
Nasim Rahaman, Devansh Arpit, Aristide Baratin, Felix Draxler, Min Lin, Fred A Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Later among the works it cites.
Implicit regularization of normalization methods
Xiaoxia Wu, Edgar Dobriban, Tongzheng Ren, Shanshan Wu, Zhiyuan Li, Suriya Gunasekar, Rachel Ward, and Qiang Liu · 2019
Later among the works it cites.
The implicit regularization of stochastic gradient flow for least squares
Alnur Ali, Edgar Dobriban, and Ryan J Tibshirani · 2020
Closest in time.
Dropout: Explicit forms and capacity control
Raman Arora, Peter Bartlett, Poorya Mianjy, and Nathan Srebro · 2020
Closest in time.
On implicit regularization: Morse functions and applications to matrix factorization
Mohamed Ali Belabbas · 2020
Closest in time.
On the inductive bias of a cnn for orthogonal patterns distributions
Alon Brutzkus and Amir Globerson · 2020
Closest in time.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Closest in time.
Can implicit bias explain generalization? stochastic convex optimization as a case study
Assaf Dauber, Meir Feder, Tomer Koren, and Roi Livni · 2020
Closest in time.
Low-rank regularization and solution uniqueness in over-parameterized matrix sensing
Kelly Geyer, Anastasios Kyrillidis, and Amir Kalev · 2020
Closest in time.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2020
Closest in time.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Closest in time.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2020
Closest in time.
Unique properties of wide minima in deep networks
Rotem Mulayoff and Tomer Michaeli · 2020
Closest in time.
Balancedness and alignment are unlikely in linear neural networks
Adityanarayanan Radhakrishnan, Eshaan Nichani, Daniel Bernstein, and Caroline Uhler · 2020
Closest in time.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Closest in time.
The implicit and explicit regularization effects of dropout
Colin Wei, Sham Kakade, and Tengyu Ma · 2020
Closest in time.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Closest in time.