Fetching the paper…
Reading the bibliography…
An important goal in reinforcement learning is to create agents that can quickly adapt to new goals while avoiding situations that might cause damage to themselves or their environments.
Safety-guided deep reinforcement learning via online gaussian process estimation
Fan, J. and Li, W. (2019) · 1903
Earlier work this paper cites.
Generalizing from a few environments in safety-critical reinforcement learning
Kenton, Z., Filos, A., Evans, O., and Gal, Y. (2019) · 1907
Earlier work this paper cites.
Safelife 1.0: Exploring side effects in complex environments
Wainwright, C. L. and Eckersley, P. (2019) · 1912
Earlier work this paper cites.
Learning to control fast-weight memories: an alternative to dynamic recurrent networks
Schmidhuber, J. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
A neural network that embeds its own meta-levels
Schmidhuber, J. (1993) · 1993
Earlier work this paper cites.
Constrained Markov decision processes
Altman, E. (1999) · 1999
Earlier work this paper cites.
Natural actor-critic
Peters, J., Vijayakumar, S., and Schaal, S. (2005) · 2005
Earlier work this paper cites.
Detection and avoidance of a carnivore odor by prey
Ferrero, D. M., Lemon, J. K., Fluegge, D., Pashkovski, S. L., Korzan, W. J., Datta, S. R., Spehr, M., Fendt, M., and Liberles, S. D. (2011) · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Safe exploration techniques for reinforcement learning–an overview
Pecka, M. and Svoboda, T. (2014) · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F. (2015) · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P. (2015) · 2015
Cited alongside, same era.
Combating deep reinforcement learning’s sisyphean curse with reinforcement learning
Lipton, Z. C., Kumar, A., Gao, J., Li, L., and Deng, L. (2016) · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
Safe reinforcement learning via shielding
Alshiekh, M., Bloem, R., Ehlers, R., Könighofer, B., Niekum, S., and Topcu, U. (2018) · 2018
Later among the works it cites.
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y. (2018) · 2018
Later among the works it cites.
Meta learning by the baldwin effect
Fernando, C. T., Sygnowski, J., Osindero, S., Wang, J., Schaul, T., Teplyashin, D., Sprechmann, P., Pritzel, A., and Rusu, A. A. (2018) · 2018
Later among the works it cites.
Pytorch implementations of reinforcement learning algorithms
Kostrikov, I. (2018) · 2018
Later among the works it cites.
Building safe artificial intelligence: specification, robustness, and assurance
Ortega, P. A., Maini, V., et al. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Itsy bitsy spider…: Infants react with increased arousal to spiders and snakes
Hoehl, S., Hellmer, K., Johansson, M., and Gredebäck, G. (2017) · 2017
Cited alongside, same era.
Deep reinforcement learning: An overview
Li, Y. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Born to learn: the inspiration, progress, and future of evolved plastic artificial neural networks
Soltoggio, A., Stanley, K. O., and Risi, S. (2017) · 2017
Cited alongside, same era.
Such, F. P., Madhavan, V., Conti, E., Lehman, J., Stanley, K. O., and Clune, J. (2017) · 2017
Cited alongside, same era.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y., Mansimov, E., Liao, S., Grosse, R., and Ba, J. (2017) · 2017
Cited alongside, same era.
Optlayer-practical constrained optimization for deep reinforcement learning in the real world
Pham, T.-H., De Magistris, G., and Tachibana, R. (2018) · 2018
Later among the works it cites.
Constrained cross-entropy method for safe reinforcement learning
Wen, M. and Topcu, U. (2018) · 2018
Later among the works it cites.
Towards continual reinforcement learning through evolutionary meta-learning
Grbic, D. and Risi, S. (2019) · 2019
Later among the works it cites.
Deep learning for video game playing
Justesen, N., Bontrager, P., Togelius, J., and Risi, S. (2019) · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D. (2019) · 2019
Later among the works it cites.
Deep neuroevolution of recurrent and discrete world models
Risi, S. and Stanley, K. O. (2019) · 2019
Later among the works it cites.