Fetching the paper…
Reading the bibliography…
We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function leads to poor performance according to the true reward function.
Understanding agent incentives using causal influence diagrams. Part I: Single action settings
Everitt, T., Ortega, P. A., Barnes, E., and Legg, S. (2019) · 1902
Earlier work this paper cites.
Conservative Agency via Attainable Utility Preservation
Turner, A. M., Hadfield-Menell, D., and Tadepalli, P. (2019) · 1902
Earlier work this paper cites.
Problems of monetary management: the UK experience
Goodhart, C. A. (1975) · 1975
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A. (1993) · 1993
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. (1999) · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al. (2000) · 2000
Earlier work this paper cites.
Glass Floor: Colleges Reject Top Applicants, Accepting Only the Students Likely to Enroll
Golden, D. (2001) · 2001
Earlier work this paper cites.
The evolved radio and its implications for modelling the evolution of novel sensors
Bird, J. and Layzell, P. (2002) · 2002
Earlier work this paper cites.
The basic AI drives
Omohundro, S. M. (2008) · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D. (2010) · 2010
Earlier work this paper cites.
Learning What to Value
Dewey, D. (2011) · 2011
Earlier work this paper cites.
Realab: An embedded perspective on tampering
Kumar, R., Uesato, J., Ngo, R., Everitt, T., Krakovna, V., and Legg, S. (2020) · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D. (2011) · 2011
Earlier work this paper cites.
Avoiding tampering incentives in deep rl via decoupled approval
Uesato, J., Kumar, R., Krakovna, V., Everitt, T., Ngo, R., and Legg, S. (2020) · 2011
Earlier work this paper cites.
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
Bostrom, N. (2012) · 2012
Earlier work this paper cites.
Brown, D. S., Schneider, J., Dragan, A. D., and Niekum, S. (2020b) · 2012
Cited alongside, same era.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N. (2014) · 2014
Cited alongside, same era.
Faulty Reward Functions in the Wild
Clark, J. and Amodei, D. (2016) · 2016
Cited alongside, same era.
Weapons of math destruction: How big data increases inequality and threatens democracy
O’Neil, C. (2016) · 2016
Cited alongside, same era.
Quantilizers: A safer alternative to maximizers for limited optimization
Taylor, J. (2016) · 2016
Cited alongside, same era.
Learning from Human Preferences
Amodei, D., Christiano, P., and Ray, A. (2017) · 2017
Cited alongside, same era.
Ambitious vs. narrow value learning
Christiano, P. (2019) · 2019
Later among the works it cites.
Artificial intelligence, values, and alignment
Gabriel, I. (2020) · 2020
Later among the works it cites.
Specification gaming: the flip side of AI ingenuity
Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., and Legg, S. (2020) · 2020
Later among the works it cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020) · 2020
Later among the works it cites.
Constrained MDPs and the reward hypothesis
Szepesvári, C. (2020) · 2020
Later among the works it cites.
Consequences of misaligned AI
Zhuang, S. and Hadfield-Menell, D. (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning with a corrupted reward channel
Everitt, T., Krakovna, V., Orseau, L., Hutter, M., and Legg, S. (2017) · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S. (2017) · 2017
Cited alongside, same era.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A. (2017) · 2017
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D. (2018) · 2018
Cited alongside, same era.
Penalizing side effects using stepwise relative reachability
Krakovna, V., Orseau, L., Kumar, R., Martic, M., and Legg, S. (2018) · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S. (2018) · 2018
Cited alongside, same era.
On the Expressivity of Markov Reward
Abel, D., Dabney, W., Harutyunyan, A., Ho, M. K., Littman, M., Precup, D., and Singh, S. (2021) · 2021
Later among the works it cites.
Hard Choices in Artificial Intelligence
Dobbe, R., Gilbert, T. K., and Mintz, Y. (2021) · 2021
Later among the works it cites.
Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
Everitt, T., Hutter, M., Kumar, R., and Krakovna, V. (2021) · 2021
Later among the works it cites.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. (2021) · 2021
Later among the works it cites.
Optimal Policies Tend to Seek Power
Turner, A. M., Smith, L., Shah, R., Critch, A., and Tadepalli, P. (2021) · 2021
Later among the works it cites.
On the Importance of Hyperparameter Optimization for Model-based Reinforcement Learning
Zhang, B., Rajan, R., Pineda, L., Lambert, N., Biedenkapp, A., Chua, K., Hutter, F., and Calandra, R. (2021) · 2021
Later among the works it cites.
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Pan, A., Bhatia, K., and Steinhardt, J. (2022) · 2022
Closest in time.
Reliance on metrics is a fundamental challenge for AI
Thomas, R. L. and Uminsky, D. (2022) · 2022
Closest in time.