Fetching the paper…
Reading the bibliography…
Learning to Optimize is a recently proposed framework for learning optimization algorithms using reinforcement learning.
Learning a synaptic learning rule
Bengio, Y, Bengio, S, and Cloutier, J · 1991
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Bengio, Yoshua · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, Sepp, Younger, A Steven, and Conwell, Peter R · 2001
Earlier work this paper cites.
A perspective view and survey of meta-learning
Vilalta, Ricardo and Drissi, Youssef · 2002
Earlier work this paper cites.
Ranking learning algorithms: Using ibl and meta-learning on accuracy and time results
Brazdil, Pavel B, Soares, Carlos, and Da Costa, Joaquim Pinto · 2003
Earlier work this paper cites.
3D hand tracking by rapid stochastic gradient descent using a skinning model
Bray, M, Koller-Meier, E, Muller, P, Van Gool, L, and Schraudolph, NN · 2004
Earlier work this paper cites.
Optimal ordered problem solver
Schmidhuber, Jürgen · 2004
Earlier work this paper cites.
Metalearning: applications to data mining
Brazdil, Pavel, Carrier, Christophe Giraud, Soares, Carlos, and Vilalta, Ricardo · 2008
Earlier work this paper cites.
Optimization on a budget: A reinforcement learning approach
Ruvolo, Paul L, Fasel, Ian, and Movellan, Javier R · 2009
Earlier work this paper cites.
Learning fast approximations of sparse coding
Gregor, Karol and LeCun, Yann · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, James S, Bardenet, Rémi, Bengio, Yoshua, and Kégl, Balázs · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2011
Cited alongside, same era.
Sequential model-based optimization for general algorithm configuration
Hutter, Frank, Hoos, Holger H, and Leyton-Brown, Kevin · 2011
Cited alongside, same era.
Random search for hyper-parameter optimization
Bergstra, James and Bengio, Yoshua · 2012
Cited alongside, same era.
Generic methods for optimization-based modeling
Domke, Justin · 2012
Cited alongside, same era.
Practical bayesian optimization of machine learning algorithms
Snoek, Jasper, Larochelle, Hugo, and Adams, Ryan P · 2012
Cited alongside, same era.
Bregman alternating direction method of multipliers
Wang, Huahua and Banerjee, Arindam · 2014
Later among the works it cites.
NIPS 1995 workshop on learning to learn: Knowledge consolidation and transfer in inductive systems
Baxter, Jonathan, Caruana, Rich, Mitchell, Tom, Pratt, Lorien Y, Silver, Daniel L, and Thrun, Sebastian · 2015
Later among the works it cites.
Initializing bayesian hyperparameter optimization via meta-learning
Feurer, Matthias, Springenberg, Jost Tobias, and Hutter, Frank · 2015
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2015
Later among the works it cites.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, Dougal, Duvenaud, David, and Adams, Ryan P · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to learn
Thrun, Sebastian and Pratt, Lorien · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
Cited alongside, same era.
Supervised sparse analysis and synthesis operators
Sprechmann, Pablo, Litman, Roee, Yakar, Tal Ben, Bronstein, Alexander M, and Sapiro, Guillermo · 2013
Cited alongside, same era.
Multi-task bayesian optimization
Swersky, Kevin, Snoek, Jasper, and Adams, Ryan P · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Later among the works it cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, Marcin, Denil, Misha, Gomez, Sergio, Hoffman, Matthew W, Pfau, David, Schaul, Tom, and de Freitas, Nando · 2016
Later among the works it cites.
Learning step size controllers for robust neural network training
Daniel, Christian, Taylor, Jonathan, and Nowozin, Sebastian · 2016
Later among the works it cites.
Deep q-networks for accelerating the training of deep neural networks
Fu, Jie, Lin, Zichuan, Liu, Miao, Leonard, Nicholas, Feng, Jiashi, and Chua, Tat-Seng · 2016
Later among the works it cites.
Using deep q-learning to control optimization hyperparameters
Hansen, Samantha · 2016
Later among the works it cites.
Li, Ke and Malik, Jitendra · 2016
Later among the works it cites.