Fetching the paper…
Reading the bibliography…
This report investigates the fitting of the Hessian or its inverse for stochastic optimizations using a Hessian fitting criterion derived from the preconditioned stochastic gradient descent (PSGD) method.
Matrix Computations
G. H. Golub and C. V. Loan · 1996
Earlier work this paper cites.
S. Amari, “Natural gradient works efficiently in learning,” Neural Computation , vol. 10, no. 2, pp. 251–276, Feb. 1998
1998
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient based learning applied to document recognition,” Proc. IEEE , vol. 86, no. 11, pp. 2278–2324, Nov. 1998
1998
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
P. Benner, E. S. Quintana-Orti and G. Quintana-Orti, “Solving stable Sylvester equations via rational iterative schemes,” Journal of Scientific Computing , vol. 28, no. 1, pp. 51–83, 2006
2006
Earlier work this paper cites.
N. J. Higham. Functions of Matrices: Theory and Computation . SIAM, 2008
2008
Earlier work this paper cites.
Journal of Machine Learning Research
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” · 2011
Earlier work this paper cites.
J. Martens and R. B. Grosse, “Optimizing neural networks with Kronecker-factored approximate curvature,” in Proc. 32nd Int. Conf. Machine Learning , 2015, pp. 2408–2417
2015
Earlier work this paper cites.
D. Povey, X. Zhang, and S. Khudanpur, “Parallel training of DNNs with natural gradient and parameter averaging,” in Proc. Int. Conf. Learning Representations , 2015
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: a method for stochastic optimization,”, in Proc. Int. Conf. Learning Representations , 2015
2015
Cited alongside, same era.
Y. N. Dauphin, H. Vries, and Y. Bengio, “Equilibrated adaptive learning rates for non-convex optimization,” in Advances in Neural Information Processing Systems , 2015, pp. 1504–1512
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: accelerating deep network training by reducing internal covariate shift,” in Proc. Int. Conf. Machine Learning , 2015
2015
Cited alongside, same era.
J. L. Ba, J. R. Kiros and G. E. Hinton, “Layer normalization,” arXiv:1607.06450, 2016
In Proc. Int. Conf. Learning Representations , 2019, New Orleans, LA, USA, May, 2019
X. L. Li, “Preconditioner on matrix Lie group for SGD,” · 2019
Later among the works it cites.
Z. Yao, A. Gholami, S. Shen, M. Mustafa, K. Keutzer, and M. W. Mahoney, “AdaHessian: an adaptive second order optimizer for machine learning,” in Proc. of the AAAI Conf. on Artificial Intelligence , 2021
2021
Later among the works it cites.
In HOOML , 2022
X. Li, “Black box Lie group preconditioners for SGD,” · 2022
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
A. Botev, H. Ritter, and D. Barber, “Practical Gauss-Newton optimisation for deep learning,” in Proc. 34nd Int. Conf. Machine Learning , 2017
2017
Cited alongside, same era.
V. Gupta, T. Koren, and Y. Singer, “Shampoo: preconditioned stochastic tensor optimization,” in Proc. 35th Int. Conf. Machine Learning , 2018
2018
Cited alongside, same era.
IEEE Trans. Neural Networks and Learning Systems
X. L. Li, “Preconditioned stochastic gradient descent,” · 2018
Cited alongside, same era.
G. Hinton, Neural Networks for Machine Learning . Retrieved from http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf
Cited in the paper.
Cited in the paper.
2023
Later among the works it cites.
S. S. Duvvuri, F. Devvrit, R. Anil, C. J. Hsieh, and I. S. Dhillon, “Combining axes preconditioners through Kronecker approximation for deep learning,” in Proc. Int. Conf. Learning Representations , 2024
2024
Closest in time.
2024
Closest in time.