Fetching the paper…
Reading the bibliography…
We consider the task of heavy-tailed statistical estimation given streaming $p$-dimensional samples.
Simple and optimal high-probability bounds for strongly-convex stochastic gradient descent
Nicholas JA Harvey, Christopher Liaw, and Sikander Randhawa · 1909
Earlier work this paper cites.
Robust regression: asymptotics, conjectures and monte carlo
Peter J Huber et al · 1973
Earlier work this paper cites.
On tail probabilities for martingales
David A Freedman · 1975
Earlier work this paper cites.
A general class of exponential inequalities for martingales and ratios
H Victor et al · 1999
Earlier work this paper cites.
An optimal median calculation algorithm for estimating internet link delays from active measurements
Dima Feldman and Yuval Shavitt · 2007
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2011
Earlier work this paper cites.
Challenging the empirical mean and empirical variance: a deviation study
Olivier Catoni · 2012
Earlier work this paper cites.
Robust lasso with missing and grossly corrupted observations
Nam H Nguyen and Trac D Tran · 2012
Earlier work this paper cites.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Earlier work this paper cites.
Efficient and fast estimation of the geometric median in hilbert spaces with an averaged stochastic gradient algorithm
Hervé Cardot, Peggy Cénac, and Pierre-André Zitt · 2013
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization, ii: shrinking procedures and optimal algorithms
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Ad click prediction: a view from the trenches
H Brendan McMahan, Gary Holt, David Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al · 2013
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2014
Earlier work this paper cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Robust regression for large-scale neuroimaging studies
Virgile Fritsch, Benoit Da Mota, Eva Loth, Gaël Varoquaux, Tobias Banaschewski, Gareth J Barker, Arun LW Bokde, Rüdiger Brühl, Brigitte Butzek, Patricia Conrod, et al · 2015
Earlier work this paper cites.
Geometric median and robust estimation in banach spaces
Stanislav Minsker et al · 2015
Cited alongside, same era.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Cited alongside, same era.
Stochastic intermediate gradient method for convex problems with stochastic inexact oracle
Pavel Dvurechensky and Alexander Gasnikov · 2016
Cited alongside, same era.
Loss minimization and parameter estimation with heavy tails
Daniel Hsu and Sivan Sabato · 2016
Cited alongside, same era.
Online estimation of the geometric median in hilbert spaces: Nonasymptotic confidence balls
Hervé Cardot, Peggy Cénac, Antoine Godichon-Baggioni, et al · 2017
Cited alongside, same era.
Being robust (in high dimensions) can be practical
Ilias Diakonikolas, Gautam Kamath, Daniel M Kane, Jerry Li, Ankur Moitra, and Alistair Stewart · 2017
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Later among the works it cites.
Algorithms of robust stochastic optimization based on mirror descent method
Alexander V Nazin, Arkadi S Nemirovsky, Alexandre B Tsybakov, and Anatoli B Juditsky · 2019
Later among the works it cites.
Adaptive hard thresholding for near-optimal consistent robust regression
Arun Sai Suggala, Kush Bhatia, Pradeep Ravikumar, and Prateek Jain · 2019
Later among the works it cites.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2019
Later among the works it cites.
Understanding gradient clipping in private sgd: A geometric perspective
Xiangyi Chen, Steven Z Wu, and Mingyi Hong · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Estimation of high dimensional mean regression in the absence of symmetry and light tail assumptions
Jianqing Fan, Quefeng Li, and Yuyan Wang · 2017
Cited alongside, same era.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron Courville · 2017
Cited alongside, same era.
Scaling sgd batch size to 32k for imagenet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
An accelerated method for derivative-free smooth stochastic convex optimization
Eduard Gorbunov, Pavel Dvurechensky, and Alexander Gasnikov · 2018
Cited alongside, same era.
Forecasting: principles and practice
Rob J Hyndman and George Athanasopoulos · 2018
Cited alongside, same era.
Fast mean estimation with sub-gaussian rates
Yeshwanth Cherapanamjeri, Nicolas Flammarion, and Peter L Bartlett · 2019
Cited alongside, same era.
High-dimensional robust mean estimation via gradient descent
Yu Cheng, Ilias Diakonikolas, Rong Ge, and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Stochastic optimization with heavy-tailed noise via accelerated gradient clipping
Eduard Gorbunov, Marina Danilova, and Alexander Gasnikov · 2020
Later among the works it cites.
Robust and heavy-tailed mean estimation made simple, via regret minimization
Sam Hopkins, Jerry Li, and Fred Zhang · 2020
Later among the works it cites.
Mean estimation with sub-gaussian rates in polynomial time
Samuel B Hopkins · 2020
Later among the works it cites.
A fast spectral algorithm for mean estimation with sub-gaussian rates
Zhixian Lei, Kyle Luh, Prayaag Venkat, and Fred Zhang · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak · 2020
Later among the works it cites.
Robust regression with covariate filtering: Heavy tails and adversarial contamination
Ankit Pensia, Varun Jog, and Po-Ling Loh · 2020
Later among the works it cites.
Robust estimation via robust gradient estimation
Adarsh Prasad, Arun Sai Suggala, Sivaraman Balakrishnan, and Pradeep Ravikumar · 2020
Later among the works it cites.
Adaptive huber regression
Qiang Sun, Wen-Xin Zhou, and Jianqing Fan · 2020
Later among the works it cites.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Later among the works it cites.
From low probability to high confidence in stochastic convex optimization
Damek Davis, Dmitriy Drusvyatskiy, Lin Xiao, and Junyu Zhang · 2021
Closest in time.
On proximal policy optimization’s heavy-tailed gradients
Saurabh Garg, Joshua Zhanson, Emilio Parisotto, Adarsh Prasad, Zico Kolter, Zachary Lipton, Sivaraman Balakrishnan, Ruslan Salakhutdinov, and Pradeep Ravikumar · 2021
Closest in time.