Fetching the paper…
Reading the bibliography…
The overall performance or expected excess risk of an iterative machine learning algorithm can be decomposed into training error and generalization error.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
A finite sample distribution-free performance bound for local discrimination rules
William H Rogers and Terry J Wagner · 1978
Earlier work this paper cites.
Distribution-free performance bounds for potential function rules
Luc P Devroye and Terry J Wagner · 1979
Earlier work this paper cites.
The jackknife, the bootstrap, and other resampling plans , volume 38
Bradley Efron · 1982
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A-S Nemirovsky, D-B Yudin, and E-R Dawson · 1982
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) \mathcal{O}(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Asymptotic methods in statistical decision theory
Lucien Le Cam · 1986
Earlier work this paper cites.
Multisurface method of pattern separation for medical diagnosis applied to breast cytology
William H Wolberg and Olvi L Mangasarian · 1990
Earlier work this paper cites.
Recursive stochastic algorithms for global optimization in ℝ d \mathbb{R}^{d}
Saul B Gelfand and Sanjoy K Mitter · 1991
Earlier work this paper cites.
Accelerated stochastic approximation
Bernard Delyon and Anatoli Juditsky · 1993
Earlier work this paper cites.
Measuring the vc-dimension of a learning machine
Vladimir Vapnik, Esther Levin, and Yann Le Cun · 1994
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Cited alongside, same era.
Almost-everywhere algorithmic stability and generalization error
Samuel Kutin and Partha Niyogi · 2002
Cited alongside, same era.
Chebyshev polynomials
John C Mason and David C Handscomb · 2002
Cited alongside, same era.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2003
Cited alongside, same era.
Boosting with the l 2 loss: regression and classification
Peter Bühlmann and Bin Yu · 2003
Cited alongside, same era.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Cited alongside, same era.
Convexity, classification, and risk bounds
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Later among the works it cites.
A lower bound for the optimization of finite sums
Alekh Agarwal and Leon Bottou · 2015
Later among the works it cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Later among the works it cites.
Global convergence of the heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2015
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Later among the works it cites.
Tight complexity bounds for optimizing composite objectives
Blake E Woodworth and Nati Srebro · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peter L Bartlett, Michael I Jordan, and Jon D McAuliffe · 2006
Cited alongside, same era.
The tradeoffs of large scale learning
Olivier Bousquet and Léon Bottou · 2008
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Cited alongside, same era.
Théorie des mécanismes connus sous le nom de parallélogrammes
Pafnuti͏̈ Lvovitch Tchebychev
Cited in the paper.
Later among the works it cites.
Stability and generalization of learning algorithms that converge to global optima
Zachary Charles and Dimitris Papailiopoulos · 2017
Later among the works it cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Later among the works it cites.
Generalization bounds of sgld for non-convex learning: Two theoretical viewpoints
Wenlong Mou, Liwei Wang, Xiyu Zhai, and Kai Zheng · 2017
Later among the works it cites.
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky · 2017
Later among the works it cites.