Fetching the paper…
Reading the bibliography…
In this paper, we study the implicit regularization of the gradient descent algorithm in homogeneous neural networks, including fully-connected and convolutional neural networks with ReLU or LeakyReLU activations.
A refined primal-dual analysis of the implicit bias
Ziwei Ji and Matus Telgarsky · 1906
Earlier work this paper cites.
The method of steepest descent for non-linear minimization problems
Haskell B Curry · 1944
Earlier work this paper cites.
Generalized gradients and applications
Frank H. Clarke · 1975
Earlier work this paper cites.
Mathematical programming methods
Guus Zoutendijk · 1976
Earlier work this paper cites.
Optimization and Nonsmooth Analysis
Frank H Clarke · 1990
Earlier work this paper cites.
Geometric categories and o-minimal structures
Lou van den Dries and Chris Miller · 1996
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E. Schapire, Yoav Freund, Peter Bartlett, and Wee Sun Lee · 1998
Earlier work this paper cites.
An Introduction to O-minimal Geometry
Michel Coste · 2002
Earlier work this paper cites.
Chapter IV - Nonsmooth Optimization Problems
Giorgio Giorgi, Angelo Guerraggio, and Jörg Thierfelder · 2004
Earlier work this paper cites.
Margin maximizing loss functions
Saharon Rosset, Ji Zhu, and Trevor J. Hastie · 2004
Earlier work this paper cites.
The dynamics of adaboost: Cyclic behavior and convergence of margins
Cynthia Rudin, Ingrid Daubechies, and Robert E Schapire · 2004
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
Pierre-Antoine Absil, Robert Mahony, and Benjamin Andrews · 2005
Earlier work this paper cites.
Analysis of boosting algorithms using the smooth margin function
Cynthia Rudin, Robert E. Schapire, and Ingrid Daubechies · 2007
Earlier work this paper cites.
Nonsmooth analysis and control theory , volume 178
Francis H. Clarke, Yuri S. Ledyaev, Ronald J. Stern, and Peter R. Wolenski · 2008
Earlier work this paper cites.
On the equivalence of weak learnability and linear separability: New relaxations and efficient boosting algorithms
Shai Shalev-Shwartz and Yoram Singer · 2010
Earlier work this paper cites.
Geometric theory of dynamical systems: an introduction
J Jr Palis and Welington De Melo · 2012
Earlier work this paper cites.
Boosting: Foundations and Algorithms
Robert E. Schapire and Yoav Freund · 2012
Earlier work this paper cites.
Evasion attacks against machine learning at test time
Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Šrndić, Pavel Laskov, Giorgio Giacinto, and Fabio Roli · 2013
Earlier work this paper cites.
Approximate KKT points and a proximity measure for termination
Joydeep Dutta, Kalyanmoy Deb, Rupesh Tulshyan, and Ramnik Arora · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Cited alongside, same era.
Margins, shrinkage, and boosting
Matus Telgarsky · 2013
Cited alongside, same era.
Curves of descent
Dmitriy Drusvyatskiy, Alexander D Ioffe, and Adrian S Lewis · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Towards a deeper geometric, analytic and algorithmic understanding of margins
Aaditya Ramdas and Javier Pena · 2016
Cited alongside, same era.
Approximation by combinations of relu and squared relu ridge functions with ℓ 1 \ell^{1} and ℓ 0 \ell^{0} controls
Jason M Klusowski and Andrew R Barron · 2018
Later among the works it cites.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2018
Later among the works it cites.
Connecting optimization and regularization paths
Arun Suggala, Adarsh Prasad, and Pradeep K Ravikumar · 2018
Later among the works it cites.
When will gradient methods converge to max-margin classifier under relu models?
Tengyu Xu, Yi Zhou, Kaiyi Ji, and Yingbin Liang · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Cited alongside, same era.
Parseval networks: Improving robustness to adversarial examples
Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier · 2017
Cited alongside, same era.
Robust large margin deep neural networks
Jure Sokolic, Raja Giryes, Guillermo Sapiro, and Miguel R. D. Rodrigues · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J. Zico Kolter, and Ryan J. Tibshirani · 2019
Closest in time.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Closest in time.
Sorting out Lipschitz function approximation
Cem Anil, James Lucas, and Roger Grosse · 2019
Closest in time.
Theory III: Dynamics and generalization in deep networks
Andrzej Banburski, Qianli Liao, Brando Miranda, Tomaso Poggio, Lorenzo Rosasco, and Jack Hidary · 2019
Closest in time.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2019
Closest in time.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Closest in time.
Implicit regularization of discrete gradient dynamics in linear neural networks
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
Bo Li, Shanshan Tang, and Haijun Yu · 2019
Closest in time.
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval, matrix completion, and blind deconvolution
Cong Ma, Kaizheng Wang, Yuejie Chi, and Yuxin Chen · 2019
Closest in time.
Regularization matters: Generalization and optimization of neural nets v.s. their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Closest in time.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 2019
Closest in time.
Stochastic subgradient method converges on tame functions
Damek Davis, Dmitriy Drusvyatskiy, Sham Kakade, and Jason D. Lee · 2020
Closest in time.
Improved sample complexities for deep neural networks and robust classification via an all-layer margin
Colin Wei and Tengyu Ma · 2020
Closest in time.