Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Yurii · 2013
Later among the works it cites.
Stochastic differential equations: an introduction with applications
Oksendal, Bernt · 2013
Later among the works it cites.
Adaptive learning rates and parallelization for stochastic, sparse, non-smooth gradients
Original
Schaul, Tom and LeCun, Yann · 2013
Later among the works it cites.
No more pesky learning rates
Schaul, Tom, Zhang, Sixin, and LeCun, Yann · 2013
Later among the works it cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Shamir, Ohad and Zhang, Tong · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
Sutskever, Ilya, Martens, James, Dahl, George, and Hinton, Geoffrey · 2013
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized algorithm
Needell, Deanna, Ward, Rachel, and Srebro, Nati · 2014
Later among the works it cites.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
Shalev-Shwartz, Shai and Zhang, Tong · 2014
Later among the works it cites.
A differential equation for modeling Nesterov’s accelerated gradient method: theory and insights
Su, Weijie, Boyd, Stephen, and Candes, Emmanuel · 2014
Later among the works it cites.
A proximal stochastic gradient method with progressive variance reduction
Xiao, Lin and Zhang, Tong · 2014
Later among the works it cites.
Probability measures for numerical solutions of differential equations
Original
Conrad, Patrick R, Girolami, Mark, Särkkä, Simo, Stuart, Andrew, and Zygalakis, Konstantinos · 2015
Closest in time.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2015
Closest in time.
Accelerated mirror descent in continuous and discrete time
Krichene, Walid, Bayen, Alexandre, and Bartlett, Peter L · 2015
Closest in time.
Continuous-time limit of stochastic gradient descent revisited
Mandt, Stephan, Hoffman, Matthew D, and Blei, David M · 2015
Closest in time.
Deep learning with elastic averaging SGD
Zhang, Sixin, Choromanska, Anna E, and LeCun, Yann · 2015
Closest in time.
Learning to learn by gradient descent by gradient descent
Andrychowicz, Marcin, Denil, Misha, Gomez, Sergio, Hoffman, Matthew W, Pfau, David, Schaul, Tom, and de Freitas, Nando · 2016
Closest in time.
A variational analysis of stochastic gradient algorithms
Original
Mandt, Stephan, Hoffman, Matthew D, and Blei, David M · 2016
Closest in time.
A variational perspective on accelerated methods in optimization
Wibisono, Andre, Wilson, Ashia C., and Jordan, Michael I · 2016
Closest in time.