Fetching the paper…
Reading the bibliography…
Preconditioned gradient methods are among the most general and powerful tools in optimization.
Über monotone matrixfunktionen
K. Löwner · 1934
Earlier work this paper cites.
Updating quasi-newton matrices with limited storage
J. Nocedal · 1980
Earlier work this paper cites.
Topics in matrix analysis, 1991
R. A. Horn and C. R. Johnson · 1991
Earlier work this paper cites.
Geometric means
T. Ando, C.-K. Li, and R. Mathias · 2004
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
N. Cesa-Bianchi, A. Conconi, and C. Gentile · 2004
Earlier work this paper cites.
Efficient algorithms for online decision problems
A. Kalai and S. Vempala · 2005
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Online learning and online convex optimization
S. Shalev-Shwartz · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2013
Cited alongside, same era.
Practical methods of optimization
R. Fletcher · 2013
Cited alongside, same era.
Nonsmooth optimization via quasi-newton methods
A. S. Lewis and M. L. Overton · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Convergence rates of sub-sampled newton methods
M. A. Erdogdu and A. Montanari · 2015
Cited alongside, same era.
Faster sgd using sketched conditioning
A. Gonen and S. Shalev-Shwartz · 2015
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Later among the works it cites.
Second order stochastic optimization in linear time
N. Agarwal, B. Bullins, and E. Hazan · 2016
Later among the works it cites.
Introduction to online convex optimization
E. Hazan · 2016
Later among the works it cites.
Sub-sampled newton methods with non-uniform sampling
P. Xu, J. Yang, F. Roosta-Khorasani, C. Ré, and M. W. Mahoney · 2016
Later among the works it cites.
A unified approach to adaptive regularization in online and stochastic optimization
V. Gupta, T. Koren, and Y. Singer · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
B. Neyshabur, R. R. Salakhutdinov, and N. Srebro · 2015
Cited alongside, same era.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
M. Pilanci and M. J. Wainwright · 2017
Later among the works it cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.