Fetching the paper…
Reading the bibliography…
We design learning rate schedules that minimize regret for SGD-based online learning in the presence of a changing data distribution.
Note on the derivatives with respect to a parameter of the solutions of a system of differential equations
T. H. Gronwall · 1919
Earlier work this paper cites.
Dynamic programming and lagrange multipliers
R. Bellman · 1956
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Varying-coefficient models
T. Hastie and R. Tibshirani · 1993
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Statistical methods with varying coefficient models
J. Fan and W. Zhang · 2008
Earlier work this paper cites.
Practical Recommendations for Gradient-Based Training of Deep Architectures , pages 437–478
Y. Bengio · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
Stochastic Differential Equations: An Introduction with Applications
B. Oksendal · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Non-stationary stochastic optimization
O. Besbes, Y. Gur, and A. Zeevi · 2015
Earlier work this paper cites.
Online optimization: Competing with dynamic comparators
A. Jadbabaie, A. Rakhlin, S. Shahrampour, and K. Sridharan · 2015
Earlier work this paper cites.
No more pesky learning rate guessing games
L. N. Smith · 2015
Earlier work this paper cites.
Deep learning with elastic averaging SGD
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Earlier work this paper cites.
TensorFlow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Cited alongside, same era.
Tracking slowly moving clairvoyant: Optimal dynamic regret of online learning with true and noisy gradient
T. Yang, L. Zhang, R. Jin, and J. Yi · 2016
Cited alongside, same era.
On the Fenchel duality between strong convexity and Lipschitz continuous gradient
X. Zhou · 2018
Later among the works it cites.
Comprehensive single cell mRNA profiling reveals a detailed roadmap for pancreatic endocrinogenesis
A. Bastidas-Ponce, L. D. Sophie Tritschler, K. Scheibner, M. Tarquis-Medina, C. Salinno, S. Schirge, I. Burtscher, A. Böttcher, F. J. Theis, H. Lickert, and M. Bakht · 2019
Later among the works it cites.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
R. Ge, S. M. Kakade, R. Kidambi, and P. Netrapalli · 2019
Later among the works it cites.
Making the last iterate of SGD information theoretically optimal
P. Jain, D. Nagaraj, and P. Netrapalli · 2019
Later among the works it cites.
Deep cytometry: Deep learning with real-time inference in cell sorting and flow cytometry
Y. Li, A. Mahjoubfar, C. L. Chen, K. R. Niazi, L. Pei, and B. Jalali · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How to escape saddle points efficiently
C. Jin, R. Ge, P. Netrapalli, S. M. Kakade, and M. I. Jordan · 2017
Cited alongside, same era.
Acceleration and averaging in stochastic descent dynamics
W. Krichene and P. L. Bartlett · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Q. Li, C. Tai, and E. Weinan · 2017
Cited alongside, same era.
Tracking moving agents via inexact online gradient descent algorithm
A. S. Bedi, P. Sarma, and K. Rajawat · 2018
Cited alongside, same era.
Deep relaxation: Partial differential equations for optimizing deep neural networks
P. Chaudhari, A. Oberman, S. Osher, S. Soatto, and G. Carlier · 2018
Cited alongside, same era.
Online bootstrap confidence intervals for the stochastic gradient descent estimator
Y. Fang, J. Xu, and L. Yang · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le · 2018
Cited alongside, same era.
An exponential learning rate schedule for deep learning
Z. Li and S. Arora · 2019
Later among the works it cites.
Generalizing rna velocity to transient cell states through dynamical modeling
V. Bergen, M. Lange, S. Peidli, F. A. Wolf, and F. Theis · 2020
Later among the works it cites.
A robust and interpretable end-to-end deep learning model for cytometry data
Z. Hu, A. Tang, J. Singh, and A. J. Butte · 2020
Later among the works it cites.
How do SGD hyperparameters in natural training affect adversarial robustness?
S. Kamath, A. Deshpande, and K. Subrahmanyam · 2020
Later among the works it cites.
On learning rates and Schrödinger operators
B. Shi, W. J. Su, and M. I. Jordan · 2020
Later among the works it cites.
On the factory floor: ML engineering for industrial-scale ads recommendation models
R. Anil, S. Gadanho, D. Huang, N. Jacob, Z. Li, D. Lin, T. Phillips, C. Pop, K. Regan, G. I. Shamir, et al · 2022
Later among the works it cites.
Application of machine learning for cytometry data
Z. Hu, S. Bhattacharya, and A. J. Butte · 2022
Later among the works it cites.
Unified Embedding: Battle-tested feature representations for web-scale ML systems
B. Coleman, W.-C. Kang, M. Fahrbach, R. Wang, L. Hong, E. H. Chi, and D. Z. Cheng · 2023
Closest in time.