Fetching the paper…
Reading the bibliography…
Despite significant progress in challenging problems across various domains, applying state-of-the-art deep reinforcement learning (RL) algorithms remains challenging due to their sensitivity to the choice of hyperparameters.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L. Lin · 1992
Earlier work this paper cites.
Evolving optimal neural networks using genetic algorithms with occam’s razor
B. Zhang and H. Mühlenbein · 1993
Earlier work this paper cites.
Genetic algorithms, tournament selection, and the effects of noise
B. L. Miller and D. E. Goldberg · 1995
Earlier work this paper cites.
A lamarckian evolution strategy for genetic algorithms
B. J Ross · 1999
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
K. O. Stanley and R. Miikkulainen · 2002
Earlier work this paper cites.
Neuroevolution: from architectures to learning
D. Floreano, P. Dürr, and C. Mattiussi · 2008
Earlier work this paper cites.
Adaptive epsilon-greedy exploration in reinforcement learning based on value difference
M. Tokic · 2010
Earlier work this paper cites.
Value-difference based exploration: Adaptive control between epsilon-greedy and softmax
M. Tokic and G. Palm · 2011
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. Adams · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
How to discount deep reinforcement learning: Towards new dynamic strategies
V. François-Lavet, R. Fonteneau, and D. Ernst · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Earlier work this paper cites.
Network morphism
T. Wei, C. Wang, Y. Rui, and C. W. Chen · 2016
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Cited alongside, same era.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
R. Islam, P. Henderson, M. Gomrokchi, and D. Precup · 2017
Cited alongside, same era.
Population based training of neural networks
M. Jaderberg, V. Dalibard, S. Osindero, W. Czarnecki, J. Donahue, A. Razavi, O. Vinyals, T. Green, I. Dunning, K. Simonyan, C. Fernando, and K. Kavukcuoglu · 2017
Cited alongside, same era.
F. P. Such, V. Madhavan, E. Conti, J. Lehman, K. Stanley, and J. Clune · 2017
Learning navigation behaviors end-to-end with autorl
H. L. Chiang, A. Faust, M. Fiser, and A. Francis · 2019
Later among the works it cites.
Evolving rewards to automate reinforcement learning
A. Faust, A. Francis, and D. Mehta · 2019
Later among the works it cites.
Hyperparameter optimization
M. Feurer and F. Hutter · 2019
Later among the works it cites.
Collaborative evolutionary reinforcement learning
S. Khadka, S. Majumdar, T. Nassar, Z. Dwiel, E. Tumer, S. Miret, Y. Liu, and K. Tumer · 2019
Later among the works it cites.
Sample-efficient deep reinforcement learning via episodic backward update
S. Y. Lee, C. Sung-Ik, and S. Chung · 2019
Later among the works it cites.
Evolving deep neural networks
R. Miikkulainen, J. Liang, E. Meyerson, A. Rawal, D. Fink, O. Francon, B. Raju, H. Shahrzad, A. Navruzyan, N. Duffy, et al · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Proceedings of the 35th International Conference on Machine Learning (ICML’18) , volume 80, 2018. Proceedings of Machine Learning Research
J. Dy and A. Krause (eds.) · 2018
Cited alongside, same era.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. G. Azar, and D. Silver · 2018
Cited alongside, same era.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. van Hasselt, and D. Silver · 2018
Cited alongside, same era.
Evolution-guided policy gradient in reinforcement learning
S. Khadka and K. Tumer · 2018
Cited alongside, same era.
Later among the works it cites.
Fast efficient hyperparameter tuning for policy gradient methods
S. Paul, V. Kurin, and S. Whiteson · 2019
Later among the works it cites.
Learning to design RNA
F. Runge, D. Stoll, S. Falkner, and F. Hutter · 2019
Later among the works it cites.
Off-policy actor-critic with shared experience replay
Simon Schmitt, Matteo Hessel, and Karen Simonyan · 2019
Later among the works it cites.
Designing neural networks through neuroevolution
K. O. Stanley, J. Clune, J. Lehman, and R. Miikkulainen · 2019
Later among the works it cites.
Proceedings of the 32nd International Conference on Advances in Neural Information Processing Systems (NeurIPS’19) , 2019
H. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alché-Buc, E. Fox, and R. Garnett (eds.) · 2019
Later among the works it cites.
Proximal distilled evolutionary reinforcement learning
C. Bodnar, B. Day, and P. Lió · 2020
Closest in time.
Provably efficient online hyperparameter optimization with population-based bandits
Jack Parker-Holder, Vu Nguyen, and Stephen J Roberts · 2020
Closest in time.
A self-tuning actor-critic algorithm
Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado P van Hasselt, David Silver, and Satinder Singh · 2020
Closest in time.