Fetching the paper…
Reading the bibliography…
In continuous action domains, standard deep reinforcement learning algorithms like DDPG suffer from inefficient exploration when facing sparse or deceptive reward problems.
Learning with Delayed Rewards
Watkins, Christopher J. C. H · 1989
Earlier work this paper cites.
Editorial
Aha, David W · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Earlier work this paper cites.
Abandoning objectives: Evolution through the search for novelty alone
Lehman, Joel and Stanley, Kenneth O · 2011
Earlier work this paper cites.
The strategic student approach for life-long exploration and learning
Lopes, Manuel and Oudeyer, Pierre-Yves · 2012
Earlier work this paper cites.
Active learning of inverse models with intrinsically motivated goal exploration in robots
Baranes, Adrien and Oudeyer, Pierre-Yves · 2013
Earlier work this paper cites.
A survey on policy search for robotics
Deisenroth, Marc Peter, Neumann, Gerhard, Peters, Jan, et al · 2013
Earlier work this paper cites.
Guided policy search
Levine, Sergey and Koltun, Vladlen · 2013
Earlier work this paper cites.
Robot skill learning: From reinforcement learning to evolution strategies
Stulp, Freek and Sigaud, Olivier · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, David, Lever, Guy, Heess, Nicolas, Degris, Thomas, Wierstra, Daan, and Riedmiller, Martin · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Earlier work this paper cites.
Confronting the challenge of quality diversity
Pugh, Justin K, Soros, LB, Szerlip, Paul A, and Stanley, Kenneth O · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, John, Levine, Sergey, Moritz, Philipp, Jordan, Michael I., and Abbeel, Pieter · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, Marc, Srinivasan, Sriram, Ostrovski, Georg, Schaul, Tom, Saxton, David, and Munos, Remi · 2016
Cited alongside, same era.
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Yan, Chen, Xi, Houthooft, Rein, Schulman, John, and Abbeel, Pieter · 2016
Cited alongside, same era.
Modular active curiosity-driven discovery of tool use
Forestier, Sébastien and Oudeyer, Pierre-Yves · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, Shixiang, Lillicrap, Timothy, Ghahramani, Zoubin, Turner, Richard E., and Levine, Sergey · 2016
Noisy networks for exploration
Fortunato, Meire, Azar, Mohammad Gheshlaghi, Piot, Bilal, Menick, Jacob, Osband, Ian, Graves, Alex, Mnih, Vlad, Munos, Remi, Hassabis, Demis, Pietquin, Olivier, et al · 2017
Later among the works it cites.
Automatic goal generation for reinforcement learning agents
Held, David, Geng, Xinyang, Florensa, Carlos, and Abbeel, Pieter · 2017
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, Peter, Islam, Riashat, Bachman, Philip, Pineau, Joelle, Precup, Doina, and Meger, David · 2017
Later among the works it cites.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Islam, Riashat, Henderson, Peter, Gomrokchi, Maziar, and Precup, Doina · 2017
Later among the works it cites.
ES is more than just a traditional finite-difference approximator
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, David, Huang, Aja, Maddison, Chris J., Guez, Arthur, Sifre, Laurent, Van Den Driessche, George, Schrittwieser, Julian, Antonoglou, Ioannis, Panneershelvam, Veda, Lanctot, Marc, et al · 2016
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, Haoran, Houthooft, Rein, Foote, Davis, Stooke, Adam, Chen, Xi, Duan, Yan, Schulman, John, De Turck, Filip, and Abbeel, Pieter · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Wang, Ziyu, Bapst, Victor, Heess, Nicolas, Mnih, Volodymyr, Munos, Remi, Kavukcuoglu, Koray, and de Freitas, Nando · 2016
Cited alongside, same era.
Bootstrapping Q-learning for robotics from neuro-evolution results
Zimmer, Matthieu and Doncieux, Stephane · 2016
Cited alongside, same era.
Andrychowicz, Marcin, Wolski, Filip, Ray, Alex, Schneider, Jonas, Fong, Rachel, Welinder, Peter, McGrew, Bob, Tobin, Josh, Abbeel, Pieter, and Zaremba, Wojciech · 2017
Cited alongside, same era.
Conti, Edoardo, Madhavan, Vashisht, Such, Felipe Petroski, Lehman, Joel, Stanley, Kenneth O., and Clune, Jeff · 2017
Cited alongside, same era.
Quality and diversity optimization: A unifying modular framework
Cully, Antoine and Demiris, Yiannis · 2017
Cited alongside, same era.
Lehman, Joel, Chen, Jay, Clune, Jeff, and Stanley, Kenneth O · 2017
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Nair, Ashvin, McGrew, Bob, Andrychowicz, Marcin, Zaremba, Wojciech, and Abbeel, Pieter · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, Deepak, Agrawal, Pulkit, Efros, Alexei A., and Darrell, Trevor · 2017
Later among the works it cites.
Petroski Such, Felipe, Madhavan, Vashisht, Conti, Edoardo, Lehman, Joel, Stanley, Kenneth O., and Clune, Jeff · 2017
Later among the works it cites.
Parameter space noise for exploration
Plappert, Matthias, Houthooft, Rein, Dhariwal, Prafulla, Sidor, Szymon, Chen, Richard Y., Chen, Xi, Asfour, Tamim, Abbeel, Pieter, and Andrychowicz, Marcin · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, Tim, Ho, Jonathan, Chen, Xi, and Sutskever, Ilya · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Yuhuai, Mansimov, Elman, Liao, Shun, Grosse, Roger, and Ba, Jimmy · 2017
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, Tuomas, Zhou, Aurick, Abbeel, Pieter, and Levine, Sergey · 2018
Closest in time.
Unsupervised learning of goal spaces for intrinsically motivated goal exploration
Pere, Alexandre, Forestier, Sebastien, Sigaud, Olivier, and Oudeyer, Pierre-Yves · 2018
Closest in time.