Fetching the paper…
Reading the bibliography…
Population risk is always of primary interest in machine learning; however, learning algorithms only have access to the empirical risk.
Evolutionsstrategie: Optimierung Technischer Systeme nach Prinzipien der Biologischen Evolution
Ingo Rechenberg and Manfred Eigen · 1973
Earlier work this paper cites.
Optimization by simulated annealing
Scott Kirkpatrick, C. D. Gelatt, and Mario Vecchi · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Exponentially many local minima for single neurons
Peter Auer, Mark Herbster, and Manfred K Warmuth · 1996
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2003
Earlier work this paper cites.
Introductory Lectures on Convex Programming
Yurii Nesterov · 2004
Earlier work this paper cites.
Online convex optimization in the bandit setting: Gradient descent without a gradient
Abraham D. Flaxman, Adam Tauman Kalai, and H. Brendan McMahan · 2005
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
Optimal algorithms for online convex optimization with multi-point bandit feedback
Alekh Agarwal, Ofer Dekel, and Lin Xiao · 2010
Earlier work this paper cites.
Differentially private empirical risk minimization
Kamalika Chaudhuri, Claire Monteleoni, and Anand D. Sarwate · 2011
Earlier work this paper cites.
Bayesian Learning via Stochastic Gradient Langevin Dynamics
Max Welling and Yee Whye Teh · 2011
Earlier work this paper cites.
Concentration Inequalities: A Nonasymptotic Theory of Independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Cited alongside, same era.
Regularized M-estimators with nonconvexity: Statistical and algorithmic theory for local optima
Po-Ling Loh and Martin J Wainwright · 2013
Cited alongside, same era.
On the complexity of bandit and derivative-free stochastic convex optimization
Ohad Shamir · 2013
Cited alongside, same era.
Escaping the Local Minima via Simulated Annealing: Optimization of Approximately Convex Functions
Alexandre Belloni, Tengyuan Liang, Hariharan Narayanan, and Alexander Rakhlin · 2015
Cited alongside, same era.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C. Duchi, Michael I. Jordan, Martin J. Wainwright, and Andre Wibisono · 2015
Cited alongside, same era.
The landscape of empirical risk for non-convex losses
Song Mei, Yu Bai, and Andrea Montanari · 2016
Later among the works it cites.
Algorithms and matching lower bounds for approximately-convex optimization
Andrej Risteski and Yuanzhi Li · 2016
Later among the works it cites.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Later among the works it cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Cited alongside, same era.
Information-theoretic lower bounds for convex optimization with erroneous oracles
Yaron Singer and Jan Vondrak · 2015
Cited alongside, same era.
Efficient approaches for escaping higher order saddle points in non-convex optimization
Animashree Anandkumar and Rong Ge · 2016
Cited alongside, same era.
Accelerated methods for non-convex optimization
Yair Carmon, John C Duchi, Oliver Hinder, and Aaron Sidford · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan
Cited in the paper.
Accelerated gradient descent escapes saddle points faster than gradient descent
Chi Jin, Praneeth Netrapalli, and Michael I. Jordan
Cited in the paper.
Certified defenses for data poisoning attacks
Jacob Steinhardt, Pang W. Koh, and Percy Liang · 2017
Later among the works it cites.
A hitting time analysis of stochastic gradient Langevin dynamics
Yuchen Zhang, Percy Liang, and Moses Charikar · 2017
Later among the works it cites.
SGD escapes saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2018
Closest in time.
An alternative view: When does SGD escape local minima?
Robert Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Closest in time.
Certifiable distributional robustness with principled adversarial training
Aman Sinha, Hongseok Namkoong, and John Duchi · 2018
Closest in time.