Fetching the paper…
Reading the bibliography…
We address the challenge of optimizing meta-parameters (hyperparameters) in machine learning, a key factor for efficient training and high model performance.
Accelerated stochastic approximation
Harry Kesten · 1958
Earlier work this paper cites.
Adaptation of learning rate parameters
Richard S. Sutton · 1981
Earlier work this paper cites.
A theory of salience change dependent on the relationship between discrepancies on successive trials on which the stimulus is present
Richard S Sutton · 1982
Earlier work this paper cites.
Increased rates of convergence through learning rate adaptation
Robert A Jacobs · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Richard S Sutton · 1992
Earlier work this paper cites.
Online independent component analysis with local learning rate adaptation
Nicol Schraudolph and Xavier Giannakopoulos · 1999
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Yoshua Bengio · 2000
Earlier work this paper cites.
Markerless tracking of complex human motions from multiple views
Roland Kehl and Luc Van Gool · 2006
Earlier work this paper cites.
Investigating Experience: Temporal Coherence and Empirical Knowledge Representation. University of Alberta MSc
A Koop · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Luke Metz, Niru Maheswaranathan, C Daniel Freeman, Ben Poole, and Jascha Sohl-Dickstein · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Neural networks for machine learning, lecture 6.5 – rmsprop, 2012
Geoffrey Hinton · 2012
Earlier work this paper cites.
Tuning-free step-size adaptation
Ashique Rupam Mahmood, Richard S Sutton, Thomas Degris, and Patrick M Pilarski · 2012
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Achieving all with no parameters: Adanormalhedge
Haipeng Luo and Robert E Schapire · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Layer-specific adaptive learning rates for deep networks
Bharat Singh, Soham De, Yangmuzi Zhang, Thomas Goldstein, and Gavin Taylor · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Cited alongside, same era.
Coin betting and parameter-free online learning
Francesco Orabona and Dávid Pál · 2016
Cited alongside, same era.
Hyperparameter optimization with approximate gradient
Fabian Pedregosa · 2016
Cited alongside, same era.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2017
Cited alongside, same era.
Training deep networks without learning rates through coin betting
Francesco Orabona and Tatiana Tommasi · 2017
Cited alongside, same era.
Bilevel programming for hyperparameter optimization and meta-learning
Gradient-based hyperparameter optimization over long horizons
Paul Micaelli and Amos J Storkey · 2021
Later among the works it cites.
Step-size adaptation using exponentiated gradient updates
Ehsan Amid, Rohan Anil, Christopher Fifty, and Manfred K Warmuth · 2022
Later among the works it cites.
Gradient descent: The ultimate optimizer
Kartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, and Erik Meijer · 2022
Later among the works it cites.
Meta mirror descent: Optimiser learning for fast convergence
Boyan Gao, Henry Gouk, Hae Beom Lee, and Timothy M Hospedales · 2022
Later among the works it cites.
Hyperparameter importance for machine learning algorithms
Honghe Jin · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
L4: Practical loss-based stepsize adaptation for deep learning
Michal Rolinek and Georg Martius · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Meta-gradient reinforcement learning
Zhongwen Xu, Hado P van Hasselt, and David Silver · 2018
Cited alongside, same era.
Metatrace: Online step-size tuning by meta-gradient descent for reinforcement learning control
Kenny Young, Baoxiang Wang, and Matthew E Taylor · 2018
Cited alongside, same era.
The importance of better models in stochastic optimization
Hilal Asi and John C Duchi · 2019
Cited alongside, same era.
Practical tradeoffs between memory, compute, and performance in learned optimizers
Luke Metz, C Daniel Freeman, James Harrison, Niru Maheswaranathan, and Jascha Sohl-Dickstein · 2022
Later among the works it cites.
A history of meta-gradient: Gradient methods for meta-learning
Richard S Sutton · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms. arxiv 2023
X Chen, C Liang, D Huang, E Real, K Wang, Y Liu, H Pham, X Dong, T Luong, CJ Hsieh, et al · 2023
Later among the works it cites.
Benchmarking neural network training algorithms
George E Dahl, Frank Schneider, Zachary Nado, Naman Agarwal, Chandramouli Shama Sastry, Philipp Hennig, Sourabh Medapati, Runa Eschenhagen, Priya Kasimbeg, Daniel Suo, et al · 2023
Later among the works it cites.
Learning-rate-free learning by d-adaptation
Aaron Defazio and Konstantin Mishchenko · 2023
Later among the works it cites.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
DoG is SGD’s best friend: A parameter-free dynamic step size schedule
Maor Ivgi, Oliver Hinder, and Yair Carmon · 2023
Later among the works it cites.
Searching for optimal per-coordinate step-sizes with multidimensional backtracking
Frederik Kunstner, Victor S Portella, Mark Schmidt, and Nick Harvey · 2023
Later among the works it cites.
Prodigy: An expeditiously adaptive parameter-free learner
Konstantin Mishchenko and Aaron Defazio · 2023
Later among the works it cites.
Mechanic: A learning rate tuner
Ashok Cutkosky, Aaron Defazio, and Harsh Mehta · 2024
Closest in time.
Step-size optimization for continual learning
Thomas Degris, Khurram Javed, Arsalan Sharifnassab, Yuxin Liu, and Richard Sutton · 2024
Closest in time.
Swifttd: A fast and robust algorithm for temporal difference learning
Khurram Javed, Arsalan Sharifnassab, and Richard S Sutton · 2024
Closest in time.
llama2.c: Inference llama 2 in one file of pure c, 2024
Andrej Karpathy · 2024
Closest in time.
Mada: Meta-adaptive optimizers through hyper-gradient descent
Kaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong, Shoham Sabach, Branislav Kveton, and Volkan Cevher · 2024
Closest in time.
Metaoptimize
Saber Salehkaleybar · 2025
Closest in time.