Fetching the paper…
Reading the bibliography…
How can we design good goals for arbitrarily intelligent agents? Reinforcement learning (RL) is a natural approach.
Anarchy, State, and Utopia
Nozick, R. (1974) · 1974
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. and Russell, S. (2000) · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
The Singularity Is Near
Kurzweil, R. (2005) · 2005
Earlier work this paper cites.
Elements of Information Theory
Cover, T. M. and Thomas, J. A. (2006) · 2006
Earlier work this paper cites.
The basic AI drives
Omohundro, S. M. (2008) · 2008
Earlier work this paper cites.
Learning what to value
Dewey, D. (2011) · 2011
Cited alongside, same era.
Delusion, survival, and intelligent agents
Ring, M. and Orseau, L. (2011) · 2011
Cited alongside, same era.
Model-based utility functions
Hibbard, B. (2012) · 2012
Cited alongside, same era.
Motivated value selection for artificial agents
Armstrong, S. (2015) · 2015
Cited alongside, same era.
Sequential extensions of causal and evidential decision theory
Everitt, T., Leike, J., and Hutter, M. (2015) · 2015
Cited alongside, same era.
Inferring human values for safe agi design
Sezener, C. E. (2015) · 2015
Cited alongside, same era.
Consequentialism
Sinnott-Armstrong, W. (2015) · 2015
The value learning problem
Soares, N. (2015) · 2015
Later among the works it cites.
Corrigibility
Soares, N., Fallenstein, B., Yudkowsky, E., and Armstrong, S. (2015) · 2015
Later among the works it cites.
Towards resolving unidentifiability in inverse reinforcement learning
Amin, K. and Singh, S. (2016) · 2016
Closest in time.
Learning the preferences of ignorant, inconsistent agents
Evans, O., Stuhlmuller, A., and Goodman, N. D. (2016) · 2016
Closest in time.
Self-modificication in rational agents
Everitt, T., Filan, D., Daswani, M., and Hutter, M. (2016) · 2016
Closest in time.
Avoiding wireheading with value reinforcement learning
Everitt, T. and Hutter, M. (2016) · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hail mary, value porosity, and utility diversification
Bostrom, N. (2014a)
Cited in the paper.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N. (2014b)
Cited in the paper.
Death and suicide in universal artificial intelligence
Martin, J., Everitt, T., and Hutter, M. (2016) · 2016
Closest in time.