Fetching the paper…
Reading the bibliography…
Deep learning has shown that learned functions can dramatically outperform hand-designed functions on perceptual tasks.
Evolutionsstrategie–optimierung technisher systeme nach prinzipien der biologischen evolution
Rechenberg, I · 1973
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o (1/kˆ 2)
Nesterov, Y · 1983
Earlier work this paper cites.
Genetic algorithms and machine learning
Goldberg, D. E. and Holland, J. H · 1988
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J. and Zipser, D · 1989
Earlier work this paper cites.
Learning a synaptic learning rule
Bengio, Y., Bengio, S., and Cloutier, J · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Werbos, P. J · 1990
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Bengio, S., Bengio, Y., Cloutier, J., and Gecsei, J · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Review papers: The statistical basis of meta-analysis
Fleiss, J · 1993
Earlier work this paper cites.
On learning how to learn learning strategies
Schmidhuber, J · 1995
Earlier work this paper cites.
No free lunch theorems for optimization
Wolpert, D. H. and Macready, W. G · 1997
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y · 1998
Earlier work this paper cites.
Gradient-based optimization of hyperparameters
Bengio, Y · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R · 2001
Earlier work this paper cites.
Natural evolution strategies
Wierstra, D., Schaul, T., Peters, J., and Schmidhuber, J · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Cited alongside, same era.
Random gradient-free minimization of convex functions
Nesterov, Y. and Spokoiny, V · 2011
Cited alongside, same era.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Cited alongside, same era.
Generic methods for optimization-based modeling
Domke, J · 2012
Cited alongside, same era.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Later among the works it cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N · 2016
Later among the works it cites.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Later among the works it cites.
Understanding short-horizon bias in stochastic meta-optimization
Wu, Y., Ren, M., Liao, R., and Grosse, R. B · 2016
Later among the works it cites.
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Staines, J. and Barber, D · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Cited alongside, same era.
Graves, A., Wayne, G., and Danihelka, I · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Bello, I., Zoph, B., Vasudevan, V., and Le, Q · 2017
Later among the works it cites.
Learning gradient descent: Better generalization and longer horizons
Lv, K., Jiang, S., and Li, J · 2017
Later among the works it cites.
Unbiasing truncated backpropagation through time
Tallec, C. and Ollivier, Y · 2017
Later among the works it cites.
Learned optimizers that scale and generalize
Wichrowska, O., Maheswaranathan, N., Hoffman, M. W., Colmenarejo, S. G., Denil, M., de Freitas, N., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Closest in time.
Bilevel programming for hyperparameter optimization and meta-learning
Franceschi, L., Frasconi, P., Salzo, S., and Pontil, M · 2018
Closest in time.
Houthooft, R., Chen, R. Y., Isola, P., Stadie, B. C., Wolski, F., Ho, J., and Abbeel, P · 2018
Closest in time.
Aggregated momentum: Stability through passive damping
Lucas, J., Zemel, R., and Grosse, R · 2018
Closest in time.
Learning unsupervised learning rules
Metz, L., Maheswaranathan, N., Cheung, B., and Sohl-Dickstein, J · 2018
Closest in time.
Pipps: Flexible model-based policy search robust to the curse of chaos
Parmas, P., Rasmussen, C. E., Peters, J., and Doya, K · 2018
Closest in time.