Fetching the paper…
Reading the bibliography…
Many applications in machine learning require optimizing a function whose true gradient is unknown, but where surrogate gradient information (directions that may be correlated with, but not necessarily identical to, the true gradient) is available instead.
About convergence of random search method in extremal control of multi-parameter systems
Rastrigin, L · 1963
Earlier work this paper cites.
Evolutionsstrategie–optimierung technisher systeme nach prinzipien der biologischen evolution
Rechenberg, I · 1973
Earlier work this paper cites.
Learning internal representations by error propagation
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1985
Earlier work this paper cites.
An efficient gradient-based algorithm for on-line training of recurrent network trajectories
Williams, R. J. and Peng, J · 1990
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Natural evolution strategies
Wierstra, D., Schaul, T., Peters, J., and Schmidhuber, J · 2008
Earlier work this paper cites.
Exponential natural evolution strategies
Glasmachers, T., Schaul, T., Yi, S., Wierstra, D., and Schmidhuber, J · 2010
Earlier work this paper cites.
Injecting external solutions into CMA-ES
Hansen, N · 2011
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Nesterov, Y. and Spokoiny, V · 2011
Earlier work this paper cites.
Generic methods for optimization-based modeling
Domke, J · 2012
Earlier work this paper cites.
Optimization for machine learning
Sra, S., Nowozin, S., and Wright, S. J · 2012
Earlier work this paper cites.
Staines, J. and Barber, D · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Random feedback weights support learning in deep neural networks
Lillicrap, T. P., Cownden, D., Tweed, D. B., and Akerman, C. J · 2014
Cited alongside, same era.
Monte carlo theory, methods and examples (book draft), 2014
Owen, A. B · 2014
Cited alongside, same era.
Optimal rates for zero-order convex optimization: The power of two function evaluations
Duchi, J. C., Jordan, M. I., Wainwright, M. J., and Wibisono, A · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Online learning rate adaptation with hypergradient descent
Baydin, A. G., Cornish, R., Rubio, D. M., Schmidt, M., and Wood, F · 2017
Later among the works it cites.
Learning to optimize
Li, K. and Malik, J · 2017
Later among the works it cites.
Learning gradient descent: Better generalization and longer horizons
Lv, K., Jiang, S., and Li, J · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Later among the works it cites.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
Tucker, G., Mnih, A., Maddison, C. J., Lawson, J., and Sohl-Dickstein, J · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N · 2016
Cited alongside, same era.
The CMA evolution strategy: A tutorial
Hansen, N · 2016
Cited alongside, same era.
Decoupled neural interfaces using synthetic gradients
Jaderberg, M., Czarnecki, W. M., Osindero, S., Vinyals, O., Graves, A., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2016
Cited alongside, same era.
Neural discrete representation learning
van den Oord, A., Vinyals, O., et al · 2017
Later among the works it cites.
Learned optimizers that scale and generalize
Wichrowska, O., Maheswaranathan, N., Hoffman, M. W., Colmenarejo, S. G., Denil, M., de Freitas, N., and Sohl-Dickstein, J · 2017
Later among the works it cites.
Structured evolution with compact architectures for scalable policy optimization
Choromanski, K., Rowland, M., Sindhwani, V., Turner, R. E., and Weller, A · 2018
Closest in time.
Evolutionary stochastic gradient descent for optimization of deep neural networks
Cui, X., Zhang, W., Tüske, Z., and Picheny, M · 2018
Closest in time.
Neuroevolution for deep reinforcement learning problems
Ha, D · 2018
Closest in time.
Ha, D. and Schmidhuber, J · 2018
Closest in time.
Houthooft, R., Chen, R. Y., Isola, P., Stadie, B. C., Wolski, F., Ho, J., and Abbeel, P · 2018
Closest in time.
Simple random search provides a competitive approach to reinforcement learning
Mania, H., Guy, A., and Recht, B · 2018
Closest in time.
CEM-RL: Combining evolutionary and gradient-based methods for policy search
Pourchot, A. and Sigaud, O · 2018
Closest in time.
Understanding short-horizon bias in stochastic meta-optimization
Wu, Y., Ren, M., Liao, R., and Grosse, R · 2018
Closest in time.