Fetching the paper…
Reading the bibliography…
We propose a framework for online meta-optimization of parameters that govern optimization, called Amortized Proximal Optimization (APO).
Fast convergence of natural gradient descent for overparameterized neural networks
Zhang, G., Martens, J., and Grosse, R · 1905
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Becker, S. and Le Cun, Y · 1988
Earlier work this paper cites.
The MNIST database of handwritten digits
LeCun, Y · 1988
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 1998
Earlier work this paper cites.
Numerical Optimization
Nocedal, J. and Wright, S · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A. and Teboulle, M · 2003
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. R · 2006
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Le Roux, N., Manzagol, P.-A., and Bengio, Y · 2007
Earlier work this paper cites.
Deep learning via Hessian-free optimization
Martens, J · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Learning recurrent neural networks with Hessian-free optimization
Martens, J. and Sutskever, I · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Bengio, Y · 2012
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
Generic methods for optimization-based modeling
Domke, J · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Lecture 6.5—RMSprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Cettolo, M., Niehues, J., Stüker, S., Bentivogli, L., and Federico, M · 2014
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
Martens, J · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Freeze-thaw Bayesian optimization
Swersky, K., Snoek, J., and Adams, R. P · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Cited alongside, same era.
Scalable Bayesian optimization using deep neural networks
Snoek, J., Rippel, O., Swersky, K., Kiros, R., Satish, N., Sundaram, N., Patwary, M., Prabhat, P., and Adams, R · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and De Freitas, N · 2016
Cited alongside, same era.
A Kronecker-factored approximate Fisher matrix for convolution layers
Grosse, R. and Martens, J · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Gradient-based meta-learning with learned layerwise metric and subspace
Lee, Y. and Choi, S · 2018
Later among the works it cites.
Kronecker-factored curvature approximations for recurrent neural networks
Martens, J., Ba, J., and Johnson, M · 2018
Later among the works it cites.
Learned optimizers that outperform SGD on wall-clock and validation loss
Metz, L., Maheswaranathan, N., Nixon, J., Freeman, C. D., and Sohl-Dickstein, J · 2018
Later among the works it cites.
Slang: Fast structured covariance approximations for Bayesian deep learning with natural gradient
Mishkin, A., Kunstner, F., Nielsen, D., Schmidt, M., and Khan, M. E · 2018
Later among the works it cites.
Approximating real-time recurrent learning with random Kronecker factors
Mujika, A., Meier, F., and Steger, A · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Li, K. and Malik, J · 2016
Cited alongside, same era.
Hyperband: Bandit-based configuration evaluation for hyperparameter optimization
Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., and Talwalkar, A · 2016
Cited alongside, same era.
Second-order optimization for neural networks. PhD Thesis
Martens, J · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Online learning rate adaptation with hypergradient descent
Baydin, A. G., Cornish, R., Rubio, D. M., Schmidt, M., and Wood, F · 2017
Cited alongside, same era.
UCI machine learning repository, 2017
Dua, D. and Graff, C · 2017
Cited alongside, same era.
Later among the works it cites.
L4: Practical loss-based stepsize adaptation for deep learning
Rolinek, M. and Martius, G · 2018
Later among the works it cites.
Understanding short-horizon bias in stochastic meta-optimization
Wu, Y., Ren, M., Liao, R., and Grosse, R · 2018
Later among the works it cites.
Stochastic (approximate) proximal point methods: Convergence, optimality, and adaptivity
Asi, H. and Duchi, J. C · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Later among the works it cites.
Optimal Kronecker-sum approximation of real time recurrent learning
Benzing, F., Gauy, M. M., Mujika, A., Martinsson, A., and Steger, A · 2019
Later among the works it cites.
MARTHE: Scheduling the learning rate via online hypergradients
Donini, M., Franceschi, L., Pontil, M., Majumder, O., and Frasconi, P · 2019
Later among the works it cites.
Meta-learning with warped gradient descent
Flennerhag, S., Rusu, A. A., Pascanu, R., Visin, F., Yin, H., and Hadsell, R · 2019
Later among the works it cites.
Revisiting the polyak step size
Hazan, E. and Kakade, S · 2019
Later among the works it cites.
Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Zhao, T · 2019
Later among the works it cites.
Aggregated momentum: Stability through passive damping
Lucas, J., Sun, S., Zemel, R., and Grosse, R · 2019
Later among the works it cites.
First-order preconditioning via hypergradient descent
Moskovitz, T., Wang, R., Lan, J., Kapoor, S., Miconi, T., Yosinski, J., and Rawal, A · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Later among the works it cites.
Park, E. and Oliva, J. B · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
Vaswani, S., Mishkin, A., Laradji, I., Schmidt, M., Gidel, G., and Lacoste-Julien, S · 2019
Later among the works it cites.
Vaswani, S., Laradji, I., Kunstner, F., Meng, S. Y., Schmidt, M., and Lacoste-Julien, S · 2020
Later among the works it cites.
University of Toronto CSC2541, Topics in Machine Learning: Neural Net Training Dynamics, Chapter 4: Second-Order Optimization
Grosse, R · 2021
Later among the works it cites.
Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence
Loizou, N., Vaswani, S., Laradji, I. H., and Lacoste-Julien, S · 2021
Later among the works it cites.
SKFAC: Training neural networks with faster Kronecker-factored approximate curvature
Tang, Z., Jiang, F., Gong, M., Li, H., Wu, Y., Yu, F., Wang, Z., and Wang, M · 2021
Later among the works it cites.
Unbiased gradient estimation in unrolled computation graphs with persistent evolution strategies
Vicol, P., Metz, L., and Sohl-Dickstein, J · 2021
Later among the works it cites.
Whitening and second order optimization both make information in the dataset unusable during training, and can reduce or prevent generalization
Wadia, N., Duckworth, D., Schoenholz, S. S., Dyer, E., and Sohl-Dickstein, J · 2021
Later among the works it cites.