Fetching the paper…
Reading the bibliography…
In this work, we present a reinforcement learning algorithm that can find a variety of policies (novel policies) for a task that is given by a task reward function.
Curious model-building control systems
Schmidhuber, J · 1991
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Outlier detection using replicator neural networks
Hawkins, S., He, H., Williams, G., and Baxter, R · 2002
Earlier work this paper cites.
Use of k-nearest neighbor classifier for intrusion detection
Liao, Y. and Vemuri, V. R · 2002
Earlier work this paper cites.
A theoretical analysis of model-based interval estimation
Strehl, A. L. and Littman, M. L · 2005
Earlier work this paper cites.
Overcoming the bootstrap problem in evolutionary robotics using behavioral diversity
Mouret, J.-B. and Doncieux, S · 2009
Earlier work this paper cites.
Anomaly detection with score functions based on nearest neighbor graphs
Zhao, M. and Saligrama, V · 2009
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Sun, Y., Gomez, F., and Schmidhuber, J · 2011
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Cited alongside, same era.
Quality diversity: A new frontier for evolutionary computation
Pugh, J. K., Soros, L. B., and Stanley, K. O · 2016
Cited alongside, same era.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Cited alongside, same era.
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y · 2017
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, O. X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Later among the works it cites.
Variational option discovery algorithms
Achiam, J., Edwards, H. A., Amodei, D., and Abbeel, P · 2018
Later among the works it cites.
Gep-pg: Decoupling exploration and exploitation in deep reinforcement learning algorithms
Colas, C., Sigaud, O., and Oudeyer, P.-Y · 2018
Later among the works it cites.
Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2017
Cited alongside, same era.
Intrinsically motivated model learning for developing curious robots
Hester, T. and Stone, P · 2017
Cited alongside, same era.
Safe visual navigation via deep learning and novelty detection
Richter, C. and Roy, N · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S
Cited in the paper.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al
Cited in the paper.
Conti, E., Madhavan, V., Such, F. P., Lehman, J., Stanley, K., and Clune, J · 2018
Later among the works it cites.
Diversity is all you need: Learning diverse skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M · 2018
Later among the works it cites.
Dart: Dynamic animation and robotics toolkit
Lee, J., Grey, M. X., Ha, S., Kunz, T., Jain, S., Ye, Y., Srinivasa, S. S., Stilman, M., and Liu, C. K · 2018
Later among the works it cites.
Deep one-class classification
Ruff, L., Görnitz, N., Deecke, L., Siddiqui, S. A., Vandermeulen, R., Binder, A., Müller, E., and Kloft, M · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.