Fetching the paper…
Reading the bibliography…
Stochastic descent methods (of the gradient and mirror varieties) have become increasingly popular in optimization.
Online learning and online convex optimization
Shai Shalev-Shwartz · 1935
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadii Nemirovskii, David Borisovich Yudin, and Edgar Ronald Dawson · 1983
Earlier work this paper cites.
A course in H-infinity control theory
Bruce A Francis · 1987
Earlier work this paper cites.
Hoo optimality criteria for LMS and backpropagation
Babak Hassibi, Ali H. Sayed, and Thomas Kailath · 1994
Earlier work this paper cites.
Hoo optimal training algorithms and their relation to backpropagation
Babak Hassibi and Thomas Kailath · 1995
Earlier work this paper cites.
Hoo optimality of the LMS algorithm
Babak Hassibi, Ali H Sayed, and Thomas Kailath · 1996
Earlier work this paper cites.
Indefinite-Quadratic Estimation and Control: A Unified Approach to H2 and H-infinity Theories , volume 16
Babak Hassibi, Ali H Sayed, and Thomas Kailath · 1999
Earlier work this paper cites.
General convergence results for linear discriminant updates
Adam J Grove, Nick Littlestone, and Dale Schuurmans · 2001
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
The robustness of the p-norm algorithms
Claudio Gentile · 2003
Earlier work this paper cites.
The p-norm generalization of the LMS algorithm for adaptive filtering
Jyrki Kivinen, Manfred K Warmuth, and Babak Hassibi · 2006
Earlier work this paper cites.
Optimal state estimation: Kalman, H infinity, and nonlinear approaches
Dan Simon · 2006
Earlier work this paper cites.
H-infinity optimal control and related minimax design problems: a dynamic game approach
Tamer Başar and Pierre Bernhard · 2008
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Mirror descent meets fixed share (and feels no regret)
Nicolo Cesa-Bianchi, Pierre Gaillard, Gábor Lugosi, and Gilles Stoltz · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Later among the works it cites.
On the emergence of invariance and disentangling in deep representations
Alessandro Achille and Stefano Soatto · 2017
Later among the works it cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Later among the works it cites.
Cong Ma, Kaizheng Wang, Yuejie Chi, and Yuxin Chen · 2017
Later among the works it cites.
Geometry of optimization and implicit regularization in deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Introduction to online convex optimization
Elad Hazan · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Cited alongside, same era.
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Later among the works it cites.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2017
Later among the works it cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Stochastic mirror descent in variationally coherent optimization problems
Zhengyuan Zhou, Panayotis Mertikopoulos, Nicholas Bambos, Stephen Boyd, and Peter W Glynn · 2017
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Closest in time.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Closest in time.