Fetching the paper…
Reading the bibliography…
Stochastic first-order methods are standard for training large-scale machine learning models.
A tail-index analysis of stochastic gradient noise in deep neural networks
Şimşekli, U., Sagun, L., and Gürbüzbalaban, M. (2019b) · 1901
Earlier work this paper cites.
On the heavy-tailed theory of stochastic gradient descent for deep neural networks
Şimşekli, U., Gürbüzbalaban, M., Nguyen, T. H., Richard, G., and Sagun, L. (2019a) · 1912
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
Probability inequalities for the sum of independent random variables
Bennett, G. (1962) · 1962
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D. A. et al. (1975) · 1975
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovsky, A. S. and Yudin, D. B. (1983) · 1983
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Nesterov, Y. E. (1983) · 1983
Earlier work this paper cites.
On bernstein-type inequalities for martingales
Dzhaparidze, K. and Van Zanten, J. (2001) · 2001
Earlier work this paper cites.
A variational formulation for frame-based inverse problems
Chaux, C., Combettes, P. L., Pesquet, J.-C., and Wajs, V. R. (2007) · 2007
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A. (2009) · 2009
Earlier work this paper cites.
First order methods for nonsmooth convex large-scale optimization, i: general purpose methods
Juditsky, A., Nemirovski, A., et al. (2011) · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E. and Bach, F. R. (2011) · 2011
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Ghadimi, S. and Lan, G. (2012) · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Lan, G. (2012) · 2012
Earlier work this paper cites.
Parametric estimation. finite sample theory
Spokoiny, V. et al. (2012) · 2012
Cited alongside, same era.
Exactness, inexactness and stochasticity in first-order methods for large-scale convex optimization
Devolder, O. (2013) · 2013
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G. (2013) · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y. (2013) · 2013
Cited alongside, same era.
First-order methods of smooth convex optimization with inexact oracle
Devolder, O., Glineur, F., and Nesterov, Y. (2014) · 2014
Cited alongside, same era.
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N. (2017) · 2017
Later among the works it cites.
Non-asymptotic confidence bounds for the optimal value of a stochastic program
Guigues, V., Juditsky, A., and Nemirovski, A. (2017) · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. (2018) · 2018
Later among the works it cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
On the efficiency of a randomized mirror descent algorithm in online optimization problems
Gasnikov, A. V., Nesterov, Y. E., and Spokoiny, V. G. (2015) · 2015
Cited alongside, same era.
On lower complexity bounds for large-scale smooth convex optimization
Guzmán, C. and Nemirovski, A. (2015) · 2015
Cited alongside, same era.
Beyond convexity: Stochastic quasi-convex optimization
Hazan, E., Levy, K., and Shalev-Shwartz, S. (2015) · 2015
Cited alongside, same era.
Universal gradient methods for convex optimization problems
Nesterov, Y. (2015) · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. (2015) · 2015
Cited alongside, same era.
Stochastic intermediate gradient method for convex problems with stochastic inexact oracle
Dvurechensky, P. and Gasnikov, A. (2016) · 2016
Cited alongside, same era.
Sgd: General analysis and improved rates
Gower, R. M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E., and Richtárik, P. (2019) · 2019
Later among the works it cites.
Algorithms of robust stochastic optimization based on mirror descent method
Nazin, A. V., Nemirovsky, A., Tsybakov, A. B., and Juditsky, A. (2019) · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. (2019) · 2019
Later among the works it cites.
Stochastic optimization with heavy-tailed noise via accelerated gradient clipping
Gorbunov, E., Danilova, M., and Gasnikov, A. (2020) · 2020
Later among the works it cites.
Can gradient clipping mitigate label noise?
Menon, A. K., Rawat, A. S., Reddi, S. J., and Kumar, S. (2020) · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. (2020) · 2020
Later among the works it cites.
From low probability to high confidence in stochastic convex optimization
Davis, D., Drusvyatskiy, D., Xiao, L., and Zhang, J. (2021) · 2021
Closest in time.
Mai, V. V. and Johansson, M. (2021) · 2021
Closest in time.