Fetching the paper…
Reading the bibliography…
Minimizing the empirical risk is a popular training strategy, but for learning tasks where the data may be noisy or heavy-tailed, one may require many observations in order to generalize well.
Handbook of Mathematical Functions With Formulas, Graphs, and Mathematical Tables
Abramowitz, M. and Stegun, I. A. (1964) · 1964
Earlier work this paper cites.
Optimization by Vector Space Methods
Luenberger, D. G. (1969) · 1969
Earlier work this paper cites.
ε \varepsilon -entropy and
Kolmogorov, A. N. (1993) · 1993
Earlier work this paper cites.
Efficient distribution-free learning of probabilistic concepts
Kearns, M. J. and Schapire, R. E. (1994) · 1994
Earlier work this paper cites.
Fat-shattering and the learnability of real-valued functions
Bartlett, P. L., Long, P. M., and Williamson, R. C. (1996) · 1996
Earlier work this paper cites.
Scale-sensitive dimensions, uniform convergence, and learnability
Alon, N., Ben-David, S., Cesa-Bianchi, N., and Haussler, D. (1997) · 1997
Earlier work this paper cites.
Asymptotic Statistics
van der Vaart, A. W. (1998) · 1998
Earlier work this paper cites.
Numerical Optimization
Nocedal, J. and Wright, S. (1999) · 1999
Earlier work this paper cites.
Probability and Measure Theory
Ash, R. B. and Doleans-Dade, C. (2000) · 2000
Earlier work this paper cites.
The multivariate
Vardi, Y. and Zhang, C.-H. (2000) · 2000
Earlier work this paper cites.
Necessary and sufficient conditions for differentiating under the integral sign
Talvila, E. (2001) · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S. (2003) · 2003
Earlier work this paper cites.
Extreme Values in Finance, Telecommunications, and the Environment
Finkenstädt, B. and Rootzén, H., editors (2003) · 2003
Cited alongside, same era.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y. (2004) · 2004
Cited alongside, same era.
High confidence estimates of the mean of heavy-tailed real random variables
Catoni, O. (2009) · 2009
Cited alongside, same era.
Robust Statistics
Huber, P. J. and Ronchetti, E. M. (2009) · 2009
Cited alongside, same era.
Robust empirical mean estimators
Lerasle, M. and Oliveira, R. I. (2011) · 2011
Cited alongside, same era.
Challenging the empirical mean and empirical variance: a deviation study
Devroye, L., Lerasle, M., Lugosi, G., and Oliveira, R. I. (2015) · 2015
Later among the works it cites.
Competing with the empirical risk minimizer in a single pass
Frostig, R., Ge, R., Kakade, S. M., and Sidford, A. (2015) · 2015
Later among the works it cites.
Geometric median and robust estimation in Banach spaces
Minsker, S. (2015) · 2015
Later among the works it cites.
Generalization of ERM in stochastic convex optimization: The dimension strikes back
Feldman, V. (2016) · 2016
Later among the works it cites.
Loss minimization and parameter estimation with heavy tails
Hsu, D. and Sabato, S. (2016) · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Catoni, O. (2012) · 2012
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for finite training sets
Le Roux, N., Schmidt, M., and Bach, F. R. (2012) · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K. (2012) · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shalev-Shwartz, S. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Optimal learners for multiclass problems
Daniely, A. and Shalev-Shwartz, S. (2014) · 2014
Cited alongside, same era.
Empirical risk minimization for heavy-tailed losses
Brownlees, C., Joly, E., and Lugosi, G. (2015) · 2015
Cited alongside, same era.
Lin, J. and Rosasco, L. (2016) · 2016
Later among the works it cites.
Risk minimization by median-of-means tournaments
Lugosi, G. and Mendelson, S. (2016) · 2016
Later among the works it cites.
Murata, T. and Suzuki, T. (2016) · 2016
Later among the works it cites.
Efficient learning with robust gradient descent
Holland, M. J. and Ikeda, K. (2017) · 2017
Closest in time.
Learning from MOM’s principles
Lecué, G. and Lerasle, M. (2017) · 2017
Closest in time.
Distributed statistical estimation and rates of convergence in normal approximation
Minsker, S. and Strawn, N. (2017) · 2017
Closest in time.
Robust estimation via robust gradient estimation
Prasad, A., Suggala, A. S., Balakrishnan, S., and Ravikumar, P. (2018) · 2018
Closest in time.