Fetching the paper…
Reading the bibliography…
We seek to align agent behavior with a user's objectives in a reinforcement learning setting with unknown dynamics, an unknown reward function, and unknown unsafe states.
The foundations of statistics
Savage, L. J · 1954
Earlier work this paper cites.
The problem of abortion and the doctrine of double effect
Foot, P · 1967
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
Premack, D. and Woodruff, G · 1978
Earlier work this paper cites.
A sequential algorithm for training text classifiers
Lewis, D. D. and Gale, W. A · 1994
Earlier work this paper cites.
The MNIST database of handwritten digits, 1998
LeCun, Y · 1998
Earlier work this paper cites.
Learning agents for uncertain environments
Russell, S. J · 1998
Earlier work this paper cites.
Active learning literature survey
Settles, B · 2009
Earlier work this paper cites.
Practical methods for optimal control and estimation using nonlinear programming
Betts, J. T · 2010
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Learning from human-generated reward
Knox, W. B · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Active reward learning
Daniel, C., Viering, M., Metz, J., Kroemer, O., and Peters, J · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F · 2015
Earlier work this paper cites.
Framing reinforcement learning from human reward: Reward positivity, temporal discounting, episodicity, and performance
Knox, W. B. and Stone, P · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Uncertainty in deep learning
Gal, Y · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A. D · 2017
Cited alongside, same era.
Active decision boundary annotation with deep generative models
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in Atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Later among the works it cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Later among the works it cites.
Mindermann, S., Shah, R., Gleave, A., and Hadfield-Menell, D · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huijser, M. and van Gemert, J. C · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Cited alongside, same era.
Interactive learning from policy-dependent human feedback
MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Wang, G., Roberts, D. L., Taylor, M. E., and Littman, M. L · 2017
Cited alongside, same era.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., and Fürnkranz, J · 2017
Cited alongside, same era.
Batch active preference-based learning of reward functions
Bıyık, E. and Sadigh, D · 2018
Cited alongside, same era.
Shared autonomy via deep reinforcement learning
Reddy, S., Dragan, A. D., and Levine, S · 2018
Later among the works it cites.
Trial without error: Towards safe reinforcement learning via human intervention
Saunders, W., Sastry, G., Stuhlmueller, A., and Evans, O · 2018
Later among the works it cites.
Active learning for convolutional neural networks: A core-set approach
Sener, O. and Savarese, S · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Deep TAMER: Interactive agent shaping in high-dimensional state spaces
Warnell, G., Waytowich, N., Lawhern, V., and Stone, P · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2019
Closest in time.
Model-based reinforcement learning for Atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2019
Closest in time.
Improving safety in reinforcement learning using model-based architectures and human intervention
Prakash, B., Khatwani, M., Waytowich, N., and Mohsenin, T · 2019
Closest in time.
The EMPATHIC framework for task learning from implicit human feedback
Cui, Y., Zhang, Q., Allievi, A., Stone, P., Niekum, S., and Knox, W. B · 2020
Closest in time.
Accelerating reinforcement learning agent with EEG-based implicit human feedback
Xu, D., Agarwal, M., Fekri, F., and Sivakumar, R · 2020
Closest in time.