Fetching the paper…
Reading the bibliography…
We propose a new stochastic gradient method called MOTAPS (Moving Targetted Polyak Stepsize) that uses recorded past loss values to compute adaptive stepsizes.
“Introduction to Optimization”
B.T. Polyak · 1987
Earlier work this paper cites.
“Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays”
U. Alon et al · 1999
Earlier work this paper cites.
“Molecular classification of cancer: class discovery and class prediction by gene expression monitoring”
T Golub et al · 1999
Earlier work this paper cites.
“Predicting the clinical status of human breast cancer by using gene expression profiles”
Mike West et al · 2001
Earlier work this paper cites.
“Online Passive-Aggressive Algorithms.”
Koby Crammer et al · 2006
Earlier work this paper cites.
“Learning multiple layers of features from tiny images”, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
“LIBSVM: a library for support vector machines”
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
“Adaptive Subgradient Methods for Online Learning and Stochastic Optimization”
John Duchi, Elad Hazan and Yoram Singer · 2011
Earlier work this paper cites.
“Non-asymptotic analysis of stochastic approximation algorithms for machine learning”
Eric Moulines and Francis Bach · 2011
Earlier work this paper cites.
“Reading Digits in Natural Images with Unsupervised Feature Learning”
Yuval Netzer et al · 2011
Earlier work this paper cites.
“Accelerating Stochastic Gradient Descent using Predictive Variance Reduction”
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
“Report on the 11th IWSLT evaluation campaign, IWSLT 2014”, 2014
Mauro Cettolo et al · 2014
Earlier work this paper cites.
“SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives”
Aaron Defazio, Francis Bach and Simon Lacoste-julien · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
“Sketch and Project: Randomized Iterative Methods for Linear Systems and Inverting Matrices”, 2016
Robert. Gower · 2016
Cited alongside, same era.
“Identity Mappings in Deep Residual Networks”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Cited alongside, same era.
“Optimizing Star-Convex Functions”
Jasper C.. Lee and Paul Valiant · 2016
Cited alongside, same era.
“UCI Machine Learning Repository”, 2017
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
“Minimizing finite sums with the stochastic average gradient”
Mark Schmidt, Nicolas Le and Francis Bach · 2017
Cited alongside, same era.
“Near-optimal methods for minimizing star-convex functions and beyond”
Oliver Hinder, Aaron Sidford and Nimit Sohoni · 2019
Later among the works it cites.
“Stochastic gradient descent with Polyak’s learning rate”
Adam. Oberman and Mariana Prazeres · 2019
Later among the works it cites.
“Unified Optimal Analysis of the (Stochastic) Gradient Method”
Sebastian. Stich · 2019
Later among the works it cites.
“Painless Stochastic Gradient: Interpolation, Line-Search, and Convergence Rates”
Sharan Vaswani et al · 2019
Later among the works it cites.
“SGD Converges to Global Minimum in Deep Learning via Star-convex Path”
Yi Zhou et al · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang et al · 2017
Cited alongside, same era.
“Gradient Descent Learns Linear Dynamical Systems”
Moritz Hardt, Tengyu Ma and Benjamin Recht · 2018
Cited alongside, same era.
“An Alternative View: When Does SGD Escape Local Minima?”
Bobby Kleinberg, Yuanzhi Li and Yang Yuan · 2018
Cited alongside, same era.
“Fast and faster convergence of SGD for over-parameterized models and an accelerated perceptron”
Sharan Vaswani, Francis Bach and Mark Schmidt · 2018
Cited alongside, same era.
“Stochastic (Approximate) Proximal Point Methods: Convergence, Optimality, and Adaptivity”
Hilal Asi and John. Duchi · 2019
Cited alongside, same era.
“Reconciling modern machine learning practice and the bias-variance trade-off”, 2019
Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal · 2019
Cited alongside, same era.
Later among the works it cites.
“Training Neural Networks for and by Interpolation”
Leonard Berrada, Andrew Zisserman and M. Kumar · 2020
Later among the works it cites.
“Better Theory for SGD in the Nonconvex World”
Ahmed Khaled and Peter Richtarik · 2020
Later among the works it cites.
“Stochastic polyak step-size for SGD: An adaptive learning rate for fast convergence”
Nicolas Loizou, Sharan Vaswani, Issam Laradji and Simon Lacoste-Julien · 2020
Later among the works it cites.
Rui Yuan, Alessandro Lazaric and Robert. Gower · 2020
Later among the works it cites.
“Statistical adaptive gradient methods”
Pengchuan Zhang, Hunger Lang, Qiang Liu and Lin Xiao · 2020
Later among the works it cites.
“SGD for Structured Nonconvex Functions: Learning Rates, Minibatching and Interpolation”
Robert. Gower, Othmane Sebbouh and Nicolas Loizou · 2021
Closest in time.
“Comment on Stochastic Polyak Step-Size: Performance of ALI-G”
M. Leonard Andrew · 2021
Closest in time.
“Almost sure convergence rates for Stochastic Gradient Descent and Stochastic Heavy Ball”
Othmane Sebbouh, Robert. Gower and Aaron Defazio · 2021
Closest in time.