Fetching the paper…
Reading the bibliography…
In this paper, we propose a new accelerated stochastic first-order method called clipped-SSTM for smooth convex stochastic optimization with heavy-tailed distributed noise in stochastic gradients and derive the first high-probability complexity bounds for this method closing the gap in the theory of stochastic optimization with heavy-tailed noise.
Cumulative frequency functions
Irving W Burr · 1942
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
A statistical distribution function of wide applicability
Waloddi Weibull · 1951
Earlier work this paper cites.
Probability inequalities for the sum of independent random variables
George Bennett · 1962
Earlier work this paper cites.
Régularisation d’inéquations variationnelles par approximations successives. rev. française informat
Bernard Martinet · 1970
Earlier work this paper cites.
Détermination approchée d’un point fixe d’une application pseudo-contractante
Bernard Martinet · 1972
Earlier work this paper cites.
On tail probabilities for martingales
David A Freedman et al · 1975
Earlier work this paper cites.
Monotone operators and the proximal point algorithm
R Tyrrell Rockafellar · 1976
Earlier work this paper cites.
Cesari convergence of the gradient method of approximating saddle points of convex-concave functions
Arkadi S Nemirovski and David Berkovich Yudin · 1978
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadi Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
On bernstein-type inequalities for martingales
Kacha Dzhaparidze and JH Van Zanten · 2001
Earlier work this paper cites.
A compendium of common probability distributions
Michael P McLaughlin · 2001
Earlier work this paper cites.
On probabilities of large deviations for random walks. i. regularly varying distribution tails
Aleksandr Alekseevich Borovkov and Konstantin Aleksandrovich Borovkov · 2002
Earlier work this paper cites.
Stability and generalization
O Bousquet A Elisseeff and Olivier Bousquet · 2002
Earlier work this paper cites.
Confidence level solutions for stochastic programming
Yu Nesterov and J-Ph Vial · 2008
Earlier work this paper cites.
On the generalization ability of online strongly convex programming algorithms
Sham M Kakade and Ambuj Tewari · 2009
Earlier work this paper cites.
Cifar-10 and cifar-100 datasets
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Stochastic convex optimization
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2009
Earlier work this paper cites.
Libsvm: A library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
Stochastic first order methods in smooth convex optimization
Olivier Devolder et al · 2011
Earlier work this paper cites.
First order methods for nonsmooth convex large-scale optimization, i: general purpose methods
Anatoli Juditsky, Arkadi Nemirovski, et al · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2011
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for svm
Shai Shalev-Shwartz, Yoram Singer, Nathan Srebro, and Andrew Cotter · 2011
Cited alongside, same era.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
Saeed Ghadimi and Guanghui Lan · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
An optimal method for stochastic composite optimization
Guanghui Lan · 2012
Cited alongside, same era.
Statistical language models based on neural networks
Tomáš Mikolov · 2012
Cited alongside, same era.
Parametric estimation. finite sample theory
Vladimir Spokoiny et al · 2012
Cited alongside, same era.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Later among the works it cites.
Robust solutions to stochastic optimization problems
Ilnura Usmanova · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Decentralize and randomize: Faster algorithm for wasserstein barycenters
Pavel Dvurechenskii, Darina Dvinskikh, Alexander Gasnikov, Cesar Uribe, and Angelia Nedich · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization, ii: shrinking procedures and optimal algorithms
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Cited alongside, same era.
Stochastic gradient methods with inexact oracle
Alexander Gasnikov, Pavel Dvurechensky, and Yurii Nesterov · 2014
Cited alongside, same era.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2014
Cited alongside, same era.
Deterministic and stochastic primal-dual subgradient algorithms for uniformly convex minimization
Anatoli Juditsky and Yuri Nesterov · 2014
Cited alongside, same era.
Eduard Gorbunov, Pavel Dvurechensky, and Alexander Gasnikov · 2018
Later among the works it cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham Kakade · 2018
Later among the works it cites.
Lectures on convex optimization
Yurii Nesterov · 2018
Later among the works it cites.
Sgd and hogwild! convergence without the bounded gradients assumption
Lam Nguyen, Phuong Ha Nguyen, Marten Dijk, Peter Richtarik, Katya Scheinberg, and Martin Takac · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Later among the works it cites.
Stochastic model-based minimization of weakly convex functions
Damek Davis and Dmitriy Drusvyatskiy · 2019
Later among the works it cites.
From low probability to high confidence in stochastic convex optimization
Damek Davis, Dmitriy Drusvyatskiy, Lin Xiao, and Junyu Zhang · 2019
Later among the works it cites.
Optimal decentralized distributed algorithms for stochastic convex optimization
Eduard Gorbunov, Darina Dvinskikh, and Alexander Gasnikov · 2019
Later among the works it cites.
A unified theory of sgd: Variance reduction, sampling, quantization and coordinate descent
Eduard Gorbunov, Filip Hanzely, and Peter Richtárik · 2019
Later among the works it cites.
Sgd: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Later among the works it cites.
A short note on concentration inequalities for random vectors with subgaussian norm
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M Kakade, and Michael I Jordan · 2019
Later among the works it cites.
Algorithms of robust stochastic optimization based on mirror descent method
Aleksandr Viktorovich Nazin, AS Nemirovsky, Aleksandr Borisovich Tsybakov, and AB Juditsky · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Later among the works it cites.
Why adam beats sgd for attention models
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2019
Later among the works it cites.
Non-monotone behavior of the heavy ball method
Marina Danilova, Anastasiia Kulakova, and Boris Polyak · 2020
Closest in time.
Better theory for sgd in the nonconvex world
Ahmed Khaled and Peter Richtárik · 2020
Closest in time.
Can gradient clipping mitigate label noise?
Aditya Krishna Menon, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2020
Closest in time.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2020
Closest in time.