Fetching the paper…
Reading the bibliography…
To improve the off-sample generalization of classical procedures minimizing the empirical risk under potentially heavy-tailed data, new robust learning algorithms have been proposed in recent years, with generalized median-of-means strategies being particularly salient.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Optimization by Vector Space Methods
Luenberger, D. G. (1969) · 1969
Earlier work this paper cites.
ε \varepsilon -entropy and
Kolmogorov, A. N. (1993) · 1993
Earlier work this paper cites.
Numerical Optimization
Nocedal, J. and Wright, S. (1999) · 1999
Earlier work this paper cites.
The multivariate
Vardi, Y. and Zhang, C.-H. (2000) · 2000
Earlier work this paper cites.
Statistical learning theory and stochastic optimization: Ecole d’Eté de Probabilités de Saint-Flour XXXI-2001
Catoni, O. (2004) · 2001
Earlier work this paper cites.
Extreme Values in Finance, Telecommunications, and the Environment
Finkenstädt, B. and Rootzén, H., editors (2003) · 2003
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y. (2004) · 2004
Earlier work this paper cites.
Challenging the empirical mean and empirical variance: a deviation study
Catoni, O. (2012) · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Optimal learners for multiclass problems
Daniely, A. and Shalev-Shwartz, S. (2014) · 2014
Cited alongside, same era.
Understanding Machine Learning: From Theory to Algorithms
Shalev-Shwartz, S. and Ben-David, S. (2014) · 2014
Cited alongside, same era.
Empirical risk minimization for heavy-tailed losses
Brownlees, C., Joly, E., and Lugosi, G. (2015) · 2015
Cited alongside, same era.
Geometric median and robust estimation in Banach spaces
Minsker, S. (2015) · 2015
Cited alongside, same era.
Loss minimization and parameter estimation with heavy tails
Hsu, D. and Sabato, S. (2016) · 2016
Later among the works it cites.
Optimal learning for multi-pass stochastic gradient methods
Lin, J. and Rosasco, L. (2016) · 2016
Later among the works it cites.
Risk minimization by median-of-means tournaments
Lugosi, G. and Mendelson, S. (2016) · 2016
Later among the works it cites.
Dimension-free PAC-Bayesian bounds for matrices, vectors, and linear least squares regression
Catoni, O. and Giulini, I. (2017) · 2017
Later among the works it cites.
Learning from MOM’s principles
Lecué, G. and Lerasle, M. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nalisnick, E., Anandkumar, A., and Smyth, P. (2015) · 2015
Cited alongside, same era.
Generalization of ERM in stochastic convex optimization: The dimension strikes back
Feldman, V. (2016) · 2016
Cited alongside, same era.
Distributed statistical machine learning in adversarial settings: Byzantine gradient descent
Chen, Y., Su, L., and Xu, J. (2017a)
Cited in the paper.
Distributed statistical machine learning in adversarial settings: Byzantine gradient descent
Chen, Y., Su, L., and Xu, J. (2017b)
Cited in the paper.
Regularization, sparse recovery, and median-of-means tournaments
Lugosi, G. and Mendelson, S. (2017a)
Cited in the paper.
Sub-gaussian estimators of the mean of a random vector
Lugosi, G. and Mendelson, S. (2017b)
Cited in the paper.
Lecué, G., Lerasle, M., and Mathieu, T. (2018) · 2018
Closest in time.
Robust estimation via robust gradient estimation
Prasad, A., Suggala, A. S., Balakrishnan, S., and Ravikumar, P. (2018) · 2018
Closest in time.