Fetching the paper…
Reading the bibliography…
Autonomous agents trained via reinforcement learning present numerous safety concerns: reward hacking, negative side effects, and unsafe exploration, among others.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J · 1991
Earlier work this paper cites.
Learning agents for uncertain environments
Russell, S · 1998
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
Knox, W. B. and Stone, P · 2009
Earlier work this paper cites.
Learning what to value
Dewey, D · 2011
Earlier work this paper cites.
April: Active preference learning-based reinforcement learning
Akrour, R., Schoenauer, M., and Sebag, M · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: a formal framework and a policy iteration algorithm
Fürnkranz, J., Hüllermeier, E., Cheng, W., and Park, S.-H · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Safe exploration techniques for reinforcement learning–an overview
Pecka, M. and Svoboda, T · 2014
Cited alongside, same era.
Corrigibility
Soares, N., Fallenstein, B., Armstrong, S., and Yudkowsky, E · 2015
Cited alongside, same era.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
Faulty reward functions in the wild , 2016
Clark, J. and Amodei, D · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A · 2016
Cited alongside, same era.
AI toy control problem , 2017
Armstrong, S · 2017
Cited alongside, same era.
Lin, Z., Harrison, B., Keech, A., and Riedl, M. O · 2017
Later among the works it cites.
Interactive learning from policy-dependent human feedback
MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Roberts, D., Taylor, M. E., and Littman, M. L · 2017
Later among the works it cites.
Dataset shift in machine learning
Sugiyama, M., Lawrence, N. D., Schwaighofer, A., et al · 2017
Later among the works it cites.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., and Fürnkranz, J · 2017
Later among the works it cites.
People teach with rewards and punishments as communication not reinforcements
Ho, M. K., Cushman, F., Littman, M. L., and Austerweil, J. L · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A · 2017
Cited alongside, same era.
Leike, J., Martic, M., Krakovna, V., Ortega, P. A., Everitt, T., Lefrancq, A., Orseau, L., and Legg, S · 2017
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Later among the works it cites.
Building safe artificial intelligence: specification, robustness, and assurance
Ortega, P. A., Maini, V., and the DeepMind safety team · 2018
Later among the works it cites.
Trial without error: Towards safe reinforcement learning via human intervention
Saunders, W., Sastry, G., Stuhlmueller, A., and Evans, O · 2018
Later among the works it cites.
Episodic curiosity through reachability
Savinov, N., Raichuk, A., Marinier, R., Vincent, D., Pollefeys, M., Lillicrap, T., and Gelly, S · 2018
Later among the works it cites.