Fetching the paper…
Reading the bibliography…
The standard L-BFGS method relies on gradient approximations that are not dominated by noise, so that search directions are descent directions, the line search is reliable, and quasi-Newton updating yields useful quadratic models of the objective function.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
Liu, D. C. and Nocedal, J · 1989
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Numerical Optimization
Nocedal, J. and Wright, S · 1999
Earlier work this paper cites.
Convex analysis and optimization
Bertsekas, D. P., Nedić, A., and Ozdaglar, A. E · 2003
Earlier work this paper cites.
Spam corpus creation for TREC
Cormack, G. and Lynam, T · 2005
Earlier work this paper cites.
A stochastic quasi-newton method for online convex optimization
Schraudolph, N. N., Yu, J., and Günter, S · 2007
Earlier work this paper cites.
Design and analysis of the causation and prediction challenge
Guyon, I., Aliferis, C. F., Cooper, G. F., Elisseeff, A., Pellet, J., Spirtes, P., and Statnikov, A. R · 2008
Earlier work this paper cites.
New probabilistic inference algorithms that harness the strengths of variational and Monte Carlo methods
Carbonetto, P · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C. and Lin, C · 2011
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
Byrd, R. H., Chin, G. M., Nocedal, J., and Wu, Y · 2012
Earlier work this paper cites.
Hybrid deterministic-stochastic methods for data fitting
Friedlander, M. P. and Schmidt, M · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Earlier work this paper cites.
Fast large-scale optimization by unifying stochastic gradient and quasi-Newton methods
Sohl-Dickstein, J., Poole, B., and Ganguli, S · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
Global convergence rate analysis of unconstrained optimization methods based on probabilistic models
Cartis, C. and Scheinberg, K · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Later among the works it cites.
adaqn: An adaptive quasi-newton algorithm for training rnns
Keskar, N. S. and Berahas, A. S · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Later among the works it cites.
A robust multi-batch l-bfgs method for machine learning
Berahas, A. S. and Takáč, M · 2017
Later among the works it cites.
Adaptive sampling strategies for stochastic optimization
Bollapragada, R., Byrd, R., and Nocedal, J · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mokhtari, A. and Ribeiro, A · 2015
Cited alongside, same era.
On sampling rates in stochastic recursions
Pasupathy, R., Glynn, P., Ghosh, S., and Hashemi, F. S · 2015
Cited alongside, same era.
Distributed second-order optimization using kronecker-factored approximations
Ba, J., Grosse, R., and Martens, J · 2016
Cited alongside, same era.
Coupling adaptive batch sizes with learning rates
Balles, L., Romero, J., and Hennig, P · 2016
Cited alongside, same era.
A multi-batch l-bfgs method for machine learning
Berahas, A. S., Nocedal, J., and Takác, M · 2016
Cited alongside, same era.
Exact and inexact subsampled Newton methods for optimization
Bollapragada, R., Byrd, R., and Nocedal, J · 2016
Cited alongside, same era.
A stochastic quasi-Newton method for large-scale optimization
Byrd, R. H., Hansen, S. L., Nocedal, J., and Singer, Y · 2016
Cited alongside, same era.
Automated inference with adaptive batches
De, S., Yadav, A., Jacobs, D., and Goldstein, T · 2017
Later among the works it cites.
Adabatch: Adaptive batch sizes for training deep neural networks
Devarakonda, A., Naumov, M., and Garland, M · 2017
Later among the works it cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Later among the works it cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Later among the works it cites.
Hoffer, E., Hubara, I., and Soudry, D · 2017
Later among the works it cites.
Neumann optimizer: A practical optimization algorithm for deep neural networks
Krishnan, S., Xiao, Y., and Saurous, R. A · 2017
Later among the works it cites.
Deep learning at 15pf: Supervised and semi-supervised classification for scientific data
Kurth, T., Zhang, J., Satish, N., Racah, E., Mitliagkas, I., Patwary, M. M. A., Malas, T., Sundaram, N., Bhimji, W., Smorkalov, M., et al · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Smith, S. L., Kindermans, P., and Le, Q. V · 2017
Later among the works it cites.
Scaling sgd batch size to 32k for imagenet training
You, Y., Gitman, I., and Ginsburg, B · 2017
Later among the works it cites.
Stochastic adaptive quasi-Newton methods for minimizing expected values
Zhou, C., Gao, W., and Goldfarb, D · 2017
Later among the works it cites.