Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning (DRL) algorithms have been successfully applied to a range of challenging control tasks.
On the theory of the brownian motion
G. E. Uhlenbeck and L. S. Ornstein · 1930
Earlier work this paper cites.
Interactions between learning and evolution
D. Ackley and M. Littman · 1991
Earlier work this paper cites.
An overview of evolutionary computation
W. M. Spears, K. A. De Jong, T. Bäck, D. B. Fogel, and H. De Garis · 1993
Earlier work this paper cites.
Evolution, learning, and instinct: 100 years of the baldwin effect
P. Turney, D. Whitley, and R. W. Anderson · 1996
Earlier work this paper cites.
Autonomous vehicle navigation using evolutionary reinforcement learning
A. Stafylopatis and K. Blekas · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
K. O. Stanley and R. Miikkulainen · 2002
Earlier work this paper cites.
Elitism-based compact genetic algorithms
C. W. Ahn and R. S. Ramakrishna · 2003
Earlier work this paper cites.
Evolutionary computation: toward a new philosophy of machine intelligence , volume 1
D. B. Fogel · 2006
Earlier work this paper cites.
Evolutionary function approximation for reinforcement learning
S. Whiteson and P. Stone · 2006
Earlier work this paper cites.
Neuroevolution: from architectures to learning
D. Floreano, P. Dürr, and C. Mattiussi · 2008
Earlier work this paper cites.
Exploiting open-endedness to solve problems through the search for novelty
J. Lehman and K. O. Stanley · 2008
Earlier work this paper cites.
Convergent temporal-difference learning with arbitrary smooth function approximation
S. Bhatnagar, D. Precup, D. Silver, R. S. Sutton, H. R. Maei, and C. Szepesvári · 2009
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Earlier work this paper cites.
Robots that can adapt like animals
A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Convolution by evolution: Differentiable pattern producing networks
C. Fernando, D. Banarse, M. Reynolds, F. Besse, D. Pfau, M. Jaderberg, M. Lanctot, and D. Wierstra · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Population based training of neural networks
M. Jaderberg, V. Dalibard, S. Osindero, W. M. Czarnecki, J. Donahue, A. Razavi, O. Vinyals, T. Green, I. Dunning, K. Simonyan, et al · 2017
Later among the works it cites.
Hierarchical representations for efficient architecture search
H. Liu, K. Simonyan, O. Vinyals, C. Fernando, and K. Kavukcuoglu · 2017
Later among the works it cites.
Continual and one-shot learning through neural networks with dynamic external memory
B. Lüders, M. Schläger, A. Korach, and S. Risi · 2017
Later among the works it cites.
Multi-step off-policy learning without importance sampling ratios
A. R. Mahmood, H. Yu, and R. S. Sutton · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Q ( λ \lambda ) with off-policy corrections
R. Munos · 2016
Cited alongside, same era.
Quality diversity: A new frontier for evolutionary computation
J. K. Pugh, L. B. Soros, and K. O. Stanley · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas · 2016
Cited alongside, same era.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
E. Conti, V. Madhavan, F. P. Such, J. Lehman, K. O. Stanley, and J. Clune · 2017
Cited alongside, same era.
G. Ostrovski, M. G. Bellemare, A. v. d. Oord, and R. Munos · 2017
Later among the works it cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Later among the works it cites.
Parameter space noise for exploration
M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz · 2017
Later among the works it cites.
Neuroevolution in games: State of the art and open challenges
S. Risi and J. Togelius · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
T. Salimans, J. Ho, X. Chen, and I. Sutskever · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
F. P. Such, V. Madhavan, E. Conti, J. Lehman, K. O. Stanley, and J. Clune · 2017
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
H. Tang, R. Houthooft, D. Foote, A. Stooke, O. X. Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel · 2017
Later among the works it cites.
Gep-pg: Decoupling exploration and exploitation in deep reinforcement learning algorithms
C. Colas, O. Sigaud, and P.-Y. Oudeyer · 2018
Closest in time.
Reinforcement learning versus evolutionary computation: A survey on hybrid algorithms
M. M. Drugan · 2018
Closest in time.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Closest in time.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Closest in time.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. van Hoof, and D. Meger · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Closest in time.
C. Sherstan, B. Bennett, K. Young, D. R. Ashley, A. White, M. White, and R. S. Sutton · 2018
Closest in time.