Fetching the paper…
Reading the bibliography…
Distribution-based search algorithms are an effective approach for evolutionary reinforcement learning of neural network controllers.
On empirical comparisons of optimizers for deep learning
Choi, D., C. J. Shallue, Z. Nado, J. Lee, C. J. Maddison, and G. E. Dahl 2019 · 1910
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T. 1964 · 1964
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
Hochreiter, S. 1991 · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. 1992 · 1992
Earlier work this paper cites.
Convergence properties of evolution strategies with the derandomized covariance matrix adaptation: The ( μ \mu / μ I \mu_{I} , λ \lambda )-cma-es
Hansen, N. and A. Ostermeier 1997 · 1997
Earlier work this paper cites.
Evolutionary algorithms and gradient search: similarities and differences
Salomon, R. 1998 · 1998
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and A. Ostermeier 2001 · 2001
Earlier work this paper cites.
Inverse mutations: Making the evolutionary-gradient-search procedure noise robust
Salomon, R. 2005 · 2005
Earlier work this paper cites.
A guide to NumPy
Oliphant, T. E. 2006 · 2006
Earlier work this paper cites.
Python 3 Reference Manual
Van Rossum, G. and F. L. Drake 2009 · 2009
Earlier work this paper cites.
Parameter-exploring policy gradients
Sehnke, F., C. Osendorfer, T. Rückstieß, A. Graves, J. Peters, and J. Schmidhuber 2010 · 2010
Earlier work this paper cites.
Infinite-horizon model predictive control for periodic tasks with contacts
Erez, T., Y. Tassa, and E. Todorov 2011 · 2011
Earlier work this paper cites.
Statistical Language Models Based on Neural Networks
Mikolov, T. 2012 · 2012
Cited alongside, same era.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Y., T. Erez, and E. Todorov 2012 · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., T. Erez, and Y. Tassa 2012 · 2012
Cited alongside, same era.
Generating sequences with recurrent neural networks
Graves, A. 2013 · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, R., T. Mikolov, and Y. Bengio 2013 · 2013
Cited alongside, same era.
Natural evolution strategies
Wierstra, D., T. Schaul, T. Glasmachers, Y. Sun, J. Peters, and J. Schmidhuber 2014 · 2014
Towards generalization and simplicity in continuous control
Rajeswaran, A., K. Lowrey, E. V. Todorov, and S. M. Kakade 2017 · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., J. Ho, X. Chen, S. Sidor, and I. Sutskever 2017 · 2017
Later among the works it cites.
Pybullet repository – issues
Coumans, E. 2018 · 2018
Later among the works it cites.
Simple random search of static linear policies is competitive for reinforcement learning
Mania, H., A. Guy, and B. Recht 2018 · 2018
Later among the works it cites.
Ray: A distributed framework for emerging AI applications
Moritz, P., R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, and I. Stoica 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and J. Ba 2015 · 2015
Cited alongside, same era.
Brockman, G., V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba 2016 · 2016
Cited alongside, same era.
The sacred infrastructure for computational research
Greff, K., A. Klein, M. Chovanec, F. Hutter, and J. Schmidhuber 2017 · 2017
Cited alongside, same era.
A visual guide to evolution strategies
Ha, D. 2017 · 2017
Cited alongside, same era.
Roboschool
Klimov, O. and J. Schulman 2017 · 2017
Cited alongside, same era.
Coumans, E. and Y. Bai 2016–2019 · 2019
Later among the works it cites.
Learning to predict without looking ahead: World models without forward prediction
Freeman, D., D. Ha, and L. Metz 2019 · 2019
Later among the works it cites.
Reinforcement learning for improving agent design
Ha, D. 2019 · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala 2019 · 2019
Later among the works it cites.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Virtanen, P., R. Gommers, T. E. Oliphant, M. Haberland, T. Reddy, D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright, S. J. van der Walt, M. Brett, J. Wilson, K. Jarrod Millman, N. Mayorov, A. R. J. Nelson, E. Jones, R. Kern, E. Larson, C. Carey, İ. Polat, Y. Feng, E. W. Moore, J. Vanderplas, D. Laxalde, J. Perktold, R. Cimrman, I. Henriksen, E. A. Quintero, C. R. Harris, A. M. Archibald, A. H. Ribeiro, F. Pedregosa, P. van Mulbregt, and Contributors 2020 · 2020
Closest in time.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Zhang, J., T. He, S. Sra, and A. Jadbabaie 2020 · 2020
Closest in time.