Fetching the paper…
Reading the bibliography…
Gradient-based hyperparameter optimization has earned a widespread popularity in the context of few-shot meta-learning, but remains broadly impractical for tasks with long horizons (many gradient steps), due to memory scaling and gradient degradation issues.
Oxford University Press, 1952
H. Stackelberg, The Theory Of Market Economy · 1952
Earlier work this paper cites.
Graduate texts in mathematics, Springer-Verlag, 1982
P. Walters, An Introduction to Ergodic Theory · 1982
Earlier work this paper cites.
“Increased rates of convergence through learning rate adaptation,” Neural Networks
1988
Earlier work this paper cites.
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” Neural Computation
1989
Earlier work this paper cites.
P. J. Werbos, “Backpropagation through time: what it does and how to do it,” Proceedings of the IEEE
1990
Earlier work this paper cites.
Y. Bengio, P. Frasconi, and P. Simard, “The problem of learning long-term dependencies in recurrent networks,” in IEEE International Conference on Neural Networks
1993
Earlier work this paper cites.
M. Riedmiller and H. Braun, “A direct adaptive method for faster backpropagation learning: the rprop algorithm,” in IEEE International Conference on Neural Networks
1993
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE Transactions on Neural Networks
1994
Earlier work this paper cites.
J. Larsen, L. K. Hansen, C. Svarer, and M. Ohlsson, “Design and regularization of neural networks: the optimal use of a validation set,” in Neural Networks for Signal Processing VI. Proceedings of the 1996 IEEE Signal Processing Society Workshop
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation
1997
Earlier work this paper cites.
Y. Bengio, “Gradient-based optimization of hyperparameters,” Neural Comput
2000
Earlier work this paper cites.
J. Domke, “Generic methods for optimization-based modeling,” in International Conference on Artificial Intelligence and Statistics
2012
Earlier work this paper cites.
R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,” in International conference on Machine Learning
2013
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International conference on Machine Learning
2015
Earlier work this paper cites.
J. Snoek, O. Rippel, K. Swersky, R. Kiros, N. Satish, N. Sundaram, M. Patwary, M. Prabhat, and R. Adams, “Scalable bayesian optimization using deep neural networks,” in International conference on Machine Learning
2015
Earlier work this paper cites.
D. Maclaurin, D. Duvenaud, and R. P. Adams, “Gradient-based Hyperparameter Optimization through Reversible Learning,” arXiv e-prints
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
F. Pedregosa, “Hyperparameter optimization with approximate gradient,” in International conference on Machine Learning
2016
Earlier work this paper cites.
J. Fu, H. Luo, J. Feng, K. H. Low, and T.-S. Chua, “Drmad: Distilling reverse-mode automatic differentiation for optimizing hyperparameters of deep neural networks,” in IJCAI
2016
Earlier work this paper cites.
J. Luketina, M. Berglund, K. Greff, and T. Raiko, “Scalable gradient-based tuning of continuous regularization hyperparameters,” in International conference on Machine Learning
2016
Cited alongside, same era.
S. Zagoruyko and N. Komodakis, “Wide residual networks,” in BMVC
2016
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Sgdr: Stochastic gradient descent with warm restarts,” in International Conference on Learning Representations
2017
Cited alongside, same era.
2017
Cited alongside, same era.
L. Li, K. G. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar, “Hyperband: A novel bandit-based approach to hyperparameter optimization,” J. Mach. Learn. Res
A. Brock, J. Donahue, and K. Simonyan, “Large scale GAN training for high fidelity natural image synthesis,” in International Conference on Learning Representations
2019
Later among the works it cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations
2019
Later among the works it cites.
M. Donini, L. Franceschi, M. Pontil, O. Majumder, and P. Frasconi, “Scheduling the Learning Rate via Hypergradients: New Insights and a New Algorithm,” arXiv e-prints
2019
Later among the works it cites.
Cham: Springer International Publishing, 2019
M. Feurer and F. Hutter, Chapter 1: Hyperparameter Optimization · 2019
Later among the works it cites.
A. Shaban, C.-A. Cheng, N. Hatch, and B. Boots, “Truncated back-propagation for bilevel optimization,” in AISTATS
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
L. Franceschi, M. Donini, P. Frasconi, and M. Pontil, “Forward and reverse gradient-based hyperparameter optimization,” in International conference on Machine Learning
2017
Cited alongside, same era.
B. Zoph and Q. V. Le, “Neural architecture search with reinforcement learning,” in International Conference on Learning Representations
2017
Cited alongside, same era.
2017
Cited alongside, same era.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
S. Falkner, A. Klein, and F. Hutter, “BOHB: Robust and efficient hyperparameter optimization at scale,” in Proceedings of the 35th International Conference on Machine Learning
2018
Cited alongside, same era.
A. Rajeswaran, C. Finn, S. M. Kakade, and S. Levine, “Meta-learning with implicit gradients,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
J. Lorraine, P. Vicol, and D. Duvenaud, “Optimizing Millions of Hyperparameters by Implicit Differentiation,” arXiv e-prints
2019
Later among the works it cites.
L. Metz, N. Maheswaranathan, J. Nixon, D. Freeman, and J. Sohl-Dickstein, “Understanding and correcting pathologies in the training of learned optimizers,” vol. 97, pp. 4556–4565, 09–15 Jun 2019
2019
Later among the works it cites.
H. Liu, K. Simonyan, and Y. Yang, “DARTS: Differentiable architecture search,” in International Conference on Learning Representations
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems
2019
Later among the works it cites.
O. Peters, “The ergodicity problem in economics,” Nature Physics
2019
Later among the works it cites.
L. Li and A. Talwalkar, “Random search and reproducibility for neural architecture search,” in UAI
2019
Later among the works it cites.
A. Klein. https://automl.github.io/HpBandSter/build/html/quickstart.html , 2019
2019
Later among the works it cites.
M. Safaryan and P. Richtárik, “On Stochastic Sign Descent Methods,” arXiv e-prints
2019
Later among the works it cites.
M. Li, E. Yumer, and D. Ramanan, “Budgeted training: Rethinking deep neural network training under resource constraints,” in International Conference on Learning Representations
2020
Closest in time.
X. Dong, M. Tan, A. W. Yu, D. Peng, B. Gabrys, and Q. V. Le, “Autohas: Differentiable hyper-parameter and architecture search,” 2020
2020
Closest in time.
2020
Closest in time.
S. Flennerhag, A. A. Rusu, R. Pascanu, F. Visin, H. Yin, and R. Hadsell, “Meta-learning with warped gradient descent,” in International Conference on Learning Representations
2020
Closest in time.
E. D. Cubuk, B. Zoph, J. Shlens, and Q. V. Le, “Randaugment: Practical automated data augmentation with a reduced search space,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
2020
Closest in time.
A. Zela, T. Elsken, T. Saikia, Y. Marrakchi, T. Brox, and F. Hutter, “Understanding and robustifying differentiable architecture search,” in International Conference on Learning Representations
2020
Closest in time.