Fetching the paper…
Reading the bibliography…
Any agent that is part of the environment it interacts with and has versatile actuators (such as arms and fingers), will in principle have the ability to self-modify -- for example by changing its own source code.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998) · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. and Barto, A. (1998) · 1998
Earlier work this paper cites.
The evolved radio and its implications for modelling the evolution of novel sensors
Bird, J. and Layzell, P. (2002) · 2002
Earlier work this paper cites.
Universal Artificial Intelligence
Hutter, M. (2005) · 2005
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
Legg, S. and Hutter, M. (2007) · 2007
Earlier work this paper cites.
Gödel machines: Fully self-referential optimal universal self-improvers
Schmidhuber, J. (2007) · 2007
Earlier work this paper cites.
The basic AI drives
Omohundro, S. M. (2008) · 2008
Earlier work this paper cites.
Learning what to value
Dewey, D. (2011) · 2011
Earlier work this paper cites.
Self-modification and mortality in artificial agents
Orseau, L. and Ring, M. (2011) · 2011
Earlier work this paper cites.
Delusion, survival, and intelligent agents
Ring, M. and Orseau, L. (2011) · 2011
Cited alongside, same era.
Model-based utility functions
Hibbard, B. (2012) · 2012
Cited alongside, same era.
Space-time embedded intelligence
Orseau, L. and Ring, M. (2012) · 2012
Cited alongside, same era.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N. (2014) · 2014
Cited alongside, same era.
Extreme state aggregation beyond MDPs
Hutter, M. (2014) · 2014
Cited alongside, same era.
General time consistent discounting
Lattimore, T. and Hutter, M. (2014) · 2014
Cited alongside, same era.
Universal knowledge-seeking agents
Orseau, L. (2014) · 2014
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., et al. (2015) · 2015
Later among the works it cites.
The value learning problem
Soares, N. (2015) · 2015
Later among the works it cites.
Corrigibility
Soares, N., Fallenstein, B., Yudkowsky, E., and Armstrong, S. (2015) · 2015
Later among the works it cites.
Artificial Superintelligence: A Futuristic Approach
Yampolskiy, R. V. (2015) · 2015
Later among the works it cites.
Self-modification of policy and utility function in rational agents
Everitt, T., Filan, D., Daswani, M., and Hutter, M. (2016) · 2016
Closest in time.
Avoiding wireheading with value reinforcement learning
Everitt, T. and Hutter, M. (2016) · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bad universal priors and notions of optimality
Leike, J. and Hutter, M. (2015) · 2015
Cited alongside, same era.
Leike, J., Lattimore, T., Orseau, L., and Hutter, M. (2016) · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., et al. (2016) · 2016
Closest in time.