Fetching the paper…
Reading the bibliography…
Training deep neural networks is a highly nontrivial task, involving carefully selecting appropriate training algorithms, scheduling step sizes and tuning other hyperparameters.
Fixed-weight networks can learn
Cotter, N.E. and Conwell, P.R · 1990
Earlier work this paper cites.
Meta-neural networks that learn by learning
Naik, D.K. and Mammone, R.J · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Learning to Learn
Thrun, S. and Pratt, L · 1998
Earlier work this paper cites.
An incremental gradient (-projection) method with momentum term and adaptive stepsize rule
Tseng, P · 1998
Earlier work this paper cites.
Simple Principles of Metalearning
Schmidhuber, J., Zhao, J., and Wiering, M · 1999
Earlier work this paper cites.
Fixed-weight on-line learning
Younger, A.S., Conwell, P.R., and Cotter, N.E · 1999
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A., and Conwell, P · 2001
Earlier work this paper cites.
Meta-learning with backpropagation
Younger, A.S., Hochreiter, S., and Conwell, P.R · 2001
Earlier work this paper cites.
Adaptive behavior with fixed weights in rnn: an overview
Prokhorov, D.V., Feldkarnp, L.A., and Tyukin, I.Y · 2002
Cited alongside, same era.
A perspective view and survey of meta-learning
Vilalta, R. and Drissi, Y · 2002
Cited alongside, same era.
Metalearning: Applications to Data Mining
Brazdil, P., Carrier, C.G., Soares, C., and Vilalta, R · 2008
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Zeiler, M.D · 2012
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Later among the works it cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G.S., Davis, A., Dean, J., Devin, M., et al · 2016
Later among the works it cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gómez, S., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N · 2016
Later among the works it cites.
Learning to learn for global optimization of black box functions
Chen, Y., Hoffman, M.W., Colmenarejo, S.G., Denil, M., Lillicrap, T.P., and de Freitas, N · 2016
Later among the works it cites.
Learning step size controllers for robust neural network training
Daniel, C., Taylor, J., and Nowozin, S · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Proximal algorithms
Parikh, N. and Boyd, S · 2014
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.A., Unterthiner, T., and Hochreiter, S · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Cited alongside, same era.
Using deep q-learning to control optimization hyperparameters
Hansen, S · 2016
Later among the works it cites.
Learning to reinforcement learn
Wang, J.X., Kurthnelson, Z., Tirumala, D., Soyer, H., Leibo, J.Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Later among the works it cites.
Learning to optimize
Li, K. and Malik, J · 2017
Closest in time.