Fetching the paper…
Reading the bibliography…
We consider the problem of how to learn a step-size policy for the Limited-Memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm.
D. C. Liu and J. Nocedal, “On the limited memory BFGS method for large scale optimization,” Mathematical programming
1989
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE
1998
Earlier work this paper cites.
Springer Science & Business Media, 1998
S. Thrun and L. Pratt, Learning to learn · 1998
Earlier work this paper cites.
Y. Bengio, “Gradient-based optimization of hyperparameters,” Neural computation
2000
Earlier work this paper cites.
S. Hochreiter, A. S. Younger, and P. R. Conwell, “Learning to learn using gradient descent,” in International Conference on Artificial Neural Networks
2001
Earlier work this paper cites.
Springer Science & Business Media, 2006
J. Nocedal and S. Wright, Numerical optimization · 2006
Earlier work this paper cites.
A. Krizhevsky, G. Hinton, et al
2009
Earlier work this paper cites.
J. S. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl, “Algorithms for hyper-parameter optimization,” in Advances in neural information processing systems
2011
Earlier work this paper cites.
J. Bergstra and Y. Bengio, “Random search for hyper-parameter optimization,” The Journal of Machine Learning Research
2012
Cited alongside, same era.
M. D. Zeiler, “Adadelta: an adaptive learning rate method,” arXiv preprint arXiv:1212.5701
2012
Cited alongside, same era.
J. Snoek, O. Rippel, K. Swersky, R. Kiros, N. Satish, N. Sundaram, M. Patwary, M. Prabhat, and R. Adams, “Scalable bayesian optimization using deep neural networks,” in International Conference on Machine Learning
2015
Cited alongside, same era.
C. Daniel, J. Taylor, and S. Nowozin, “Learning step size controllers for robust neural network training.,” in AAAI
2016
Cited alongside, same era.
K. Li and J. Malik, “Learning to optimize,” arXiv preprint arXiv:1606.01885
C. Zhou, W. Gao, and D. Goldfarb, “Stochastic adaptive quasi-newton methods for minimizing expected values,” in International Conference on Machine Learning
2017
Later among the works it cites.
X. Dong, J. Shen, W. Wang, Y. Liu, L. Shao, and F. Porikli, “Hyperparameter optimization for tracking with continuous deep Q-learning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition
2018
Later among the works it cites.
2018
Later among the works it cites.
L. Metz, N. Maheswaranathan, J. Nixon, D. Freeman, and J. Sohl-Dickstein, “Understanding and correcting pathologies in the training of learned optimizers,” in International Conference on Machine Learning
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
P. Moritz, R. Nishihara, and M. Jordan, “A linearly-convergent stochastic L-BFGS algorithm,” in Artificial Intelligence and Statistics
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Agrawal, B. Amos, S. Barratt, S. Boyd, S. Diamond, and Z. Kolter, “Differentiable convex optimization layers,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, et al
2019
Later among the works it cites.