Fetching the paper…
Reading the bibliography…
Optimization in machine learning, both theoretical and applied, is presently dominated by first-order gradient methods such as stochastic gradient descent.
Updating quasi-newton matrices with limited storage
J. Nocedal · 1980
Earlier work this paper cites.
Trust region methods
A. R. Conn, N. I. Gould, and P. L. Toint · 2000
Earlier work this paper cites.
On “natural” learning and pruning in multilayered perceptrons
T. Heskes · 2000
Earlier work this paper cites.
Geometric means
T. Ando, C.-K. Li, and R. Mathias · 2004
Earlier work this paper cites.
A Schur-Newton method for the matrix p’th root and its inverse
C.-H. Guo and N. J. Higham · 2006
Earlier work this paper cites.
On the Newton method for the matrix p-th root
B. Iannazzo · 2006
Earlier work this paper cites.
Numerical optimization
J. Nocedal and S. Wright · 2006
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
E. Hazan, A. Agarwal, and S. Kale · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky et al · 2009
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. B. McMahan and M. Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. A. Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng · 2012
Earlier work this paper cites.
Japanese and Korean voice search
M. Schuster and K. Nakajima · 2012
Earlier work this paper cites.
Online learning and online convex optimization
S. Shalev-Shwartz · 2012
Earlier work this paper cites.
Practical methods of optimization
R. Fletcher · 2013
Earlier work this paper cites.
Nonsmooth optimization via quasi-newton methods
A. S. Lewis and M. L. Overton · 2013
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, J. Leveling, C. Monz, P. Pecina, M. Post, H. Saint-Amand, R. Soricut, L. Specia, and A. s. Tamchyna · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Criteo releases industry’s largest-ever dataset for machine learning to academic community, July 2015
Criteo Labs · 2015
Earlier work this paper cites.
Convergence rates of sub-sampled newton methods
M. A. Erdogdu and A. Montanari · 2015
Earlier work this paper cites.
Faster sgd using sketched conditioning
A. Gonen and S. Shalev-Shwartz · 2015
Cited alongside, same era.
Deep learning with limited numerical precision
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan · 2015
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Cited alongside, same era.
Shampoo: Preconditioned stochastic tensor optimization
V. Gupta, T. Koren, and Y. Singer · 2018
Later among the works it cites.
Mesh-tensorflow: Deep learning for supercomputers
N. Shazeer, Y. Cheng, N. Parmar, D. Tran, A. Vaswani, P. Koanantakool, P. Hawkins, H. Lee, M. Hong, C. Young, et al · 2018
Later among the works it cites.
Leveraging the bfloat16 artificial intelligence datatype for higher-precision computations
G. Henry, P. T. P. Tang, and A. Heinecke · 2019
Later among the works it cites.
Simulating low precision floating-point arithmetic
N. J. Higham and S. Pranesh · 2019
Later among the works it cites.
Limitations of the empirical fisher approximation for natural gradient descent
F. Kunstner, P. Hennig, and L. Balles · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Agarwal, B. Bullins, and E. Hazan · 2016
Cited alongside, same era.
Introduction to online convex optimization
E. Hazan · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Sub-sampled newton methods with non-uniform sampling
P. Xu, J. Yang, F. Roosta-Khorasani, C. Ré, and M. W. Mahoney · 2016
Cited alongside, same era.
Distributed second-order optimization using kronecker-factored approximations
J. Ba, J. Martens, and R. Grosse · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, et al · 2017
Cited alongside, same era.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
M. Pilanci and M. J. Wainwright · 2017
Cited alongside, same era.
P. Mattson, C. Cheng, C. Coleman, G. Diamos, P. Micikevicius, D. Patterson, H. Tang, G.-Y. Wei, P. Bailis, V. Bittorf, et al · 2019
Later among the works it cites.
Deep learning recommendation model for personalization and recommendation systems
M. Naumov, D. Mudigere, H.-J. M. Shi, J. Huang, N. Sundaraman, J. Park, X. Wang, U. Gupta, C.-J. Wu, A. G. Azzolini, et al · 2019
Later among the works it cites.
Large-scale distributed second-order optimization using kronecker-factored approximate curvature for deep convolutional neural networks
K. Osawa, Y. Tsuji, Y. Ueno, A. Naruse, R. Yokota, and S. Matsuoka · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Later among the works it cites.
Lingvo: a modular and scalable framework for sequence-to-sequence modeling, 2019
J. Shen, P. Nguyen, Y. Wu, Z. Chen, et al · 2019
Later among the works it cites.
Bfloat16: The secret to high performance on cloud tpus
S. Wang and P. Kanwar · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh · 2019
Later among the works it cites.
Disentangling adaptive gradient methods from learning rates
N. Agarwal, R. Anil, E. Hazan, T. Koren, and C. Zhang · 2020
Closest in time.
Practical quasi-newton methods for training deep neural networks
D. Goldfarb, Y. Ren, and A. Bahamou · 2020
Closest in time.
Training v0.7 results
MLPerf · 2020
Closest in time.
Distributed equivalent substitution training for large-scale recommender systems
H. Rong, Y. Wang, F. Zhou, J. Zhai, H. Wu, R. Lan, F. Li, H. Zhang, Y. Yang, Z. Guo, et al · 2020
Closest in time.
Developing a recommendation benchmark for mlperf training and inference
C.-J. Wu, R. Burke, E. Chi, J. Konstan, J. McAuley, Y. Raimond, and H. Zhang · 2020
Closest in time.
Automatic cross-replica sharding of weight update in data-parallel training
Y. Xu, H. Lee, D. Chen, H. Choi, B. Hechtman, and S. Wang · 2020
Closest in time.
A large batch optimizer reality check: Traditional, generic optimizers suffice across batch sizes
Z. Nado, J. M. Gilmer, C. J. Shallue, R. Anil, and G. E. Dahl · 2021
Closest in time.