Fetching the paper…
Reading the bibliography…
We establish novel generalization bounds for learning algorithms that converge to global minima.
A topological property of real analytic subsets
S Lojasiewicz · 1963
Earlier work this paper cites.
Distribution-free performance bounds for potential function rules
L. Devroye and T. Wagner · 1979
Earlier work this paper cites.
On sensitivity analysis of nonlinear programs in banach spaces: the approach via composite unconstrained optimization
Alexander Ioffe · 1994
Earlier work this paper cites.
Second-order sufficiency and quadratic growth for nonisolated minima
Joseph Frédéric Bonnans and Alexander Ioffe · 1995
Earlier work this paper cites.
Degenerate nonlinear programming with a quadratic growth condition
Mihai Anitescu · 2000
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Stability of randomized learning algorithms
Andre Elisseeff, Theodoros Evgeniou, and Massimiliano Pontil · 2005
Earlier work this paper cites.
Differential privacy
Cynthia Dwork · 2006
Earlier work this paper cites.
Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
Sayan Mukherjee, Partha Niyogi, Tomaso Poggio, and Ryan Rifkin · 2006
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Yu Nesterov · 2012
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2013
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Preserving statistical validity in adaptive data analysis
Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Leon Roth · 2015
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Later among the works it cites.
Why does deep and cheap learning work so well?
Henry W Lin and Max Tegmark · 2016
Later among the works it cites.
Generalization properties and implicit regularization for multiple passes sgm
Junhong Lin, Raffaello Camoriano, and Lorenzo Rosasco · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the generalization properties of differential privacy
Kobbi Nissim and Uri Stemmer · 2015
Cited alongside, same era.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Cited alongside, same era.
Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak-Łojasiewicz Condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Fast rates for empirical risk minimization of strict saddle problems
Alon Gonen and Shai Shalev-Shwartz · 2017
Closest in time.
Data-dependent stability of stochastic gradient descent
Ilja Kuzborskij and Christoph Lampert · 2017
Closest in time.
Algorithmic stability and hypothesis complexity
Tongliang Liu, Gábor Lugosi, Gergely Neu, and Dacheng Tao · 2017
Closest in time.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2017
Closest in time.
The landscape of deep learning algorithms
Pan Zhou and Jiashi Feng · 2017
Closest in time.