Fetching the paper…
Reading the bibliography…
We present a novel approach to accelerate stochastic gradient descent (SGD) by utilizing curvature information obtained from Hessian-vector products or finite differences of parameters and gradients, similar to the BFGS algorithm.
Yuan Cao and Quanquan Gu · 1902
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz · 2004
Earlier work this paper cites.
Practical quasi-newton methods for training deep neural networks
Donald Goldfarb, Yi Ren, and Achraf Bahamou · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Towards theoretically understanding why sgd generalizes better than adam in deep learning, 2020a
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven Hoi, and Weinan E · 2010
Earlier work this paper cites.
Training deep and recurrent neural networks with hessian-free optimization
J. Martens and I. Sutskever · 2012
Earlier work this paper cites.
Equilibrated adaptive learning rates for non-convex optimization
Y. N. Dauphin, H. Vries, and Y. Bengio · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: a method for stochastic optimization
D. P. Kingma and J. L. Ba · 2015
Earlier work this paper cites.
Preconditioned stochastic gradient descent, 2015
Xi-Lin Li · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
J. Martens and R. B. Grosse · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
An empirical study of example forgetting during deep neural network learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon · 2018
Later among the works it cites.
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Later among the works it cites.
Openwebtext corpus, 2019
Aaron Gokaslan and Vanya Cohen · 2019
Later among the works it cites.
Preconditioner on matrix lie group for SGD
X. L. Li · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Second-order Optimization for Neural Networks
James Martens · 2016
Cited alongside, same era.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Residual networks behave like ensembles of relatively shallow networks, 2016
Andreas Veit, Michael Wilber, and Serge Belongie · 2016
Cited alongside, same era.
Critical learning periods in deep neural networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V. Ugur Güney, Yann N. Dauphin, and Léon Bottou · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
torch-optimizer – collection of optimization algorithms for PyTorch., January 2020
Mykola Novik · 2020
Later among the works it cites.
Adabelief optimizer: adapting stepsizes by the belief in observed gradients
Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James S. Duncan · 2020
Later among the works it cites.
Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
Jungmin Kwon, Jeongseop Kim, Hyunseo Park, and In Kwon Choi · 2021
Later among the works it cites.
Adahessian: an adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael W. Mahoney · 2021
Later among the works it cites.
Neural tangent generalization attacks
Chia-Hung Yuan and Shan-Hung Wu · 2021
Later among the works it cites.
Asdl: A unified interface for gradient preconditioning in pytorch
Kazuki Osawa, Satoki Ishikawa, Rio Yokota, Shigang Li, and Torsten Hoefler · 2022
Later among the works it cites.
Adaptive second order coresets for data-efficient machine learning
Omead Pooladzandi, David Davini, and Baharan Mirzasoleiman · 2022
Later among the works it cites.
Deep learning without shortcuts: Shaping the kernel with tailored rectifiers
Guodong Zhang, Aleksandar Botev, and James Martens · 2022
Later among the works it cites.
Nanogpt: Small gpt implementations
Andrej Karpathy · 2023
Later among the works it cites.
Generating high fidelity synthetic data via coreset selection and entropic regularization, 2023
Omead Pooladzandi, Pasha Khosravi, Erik Nijkamp, and Baharan Mirzasoleiman · 2023
Later among the works it cites.