Fetching the paper…
Reading the bibliography…
Implementing a reward function that perfectly captures a complex task in the real world is impractical.
The steepest-ascent method for the linear programming problem
J. Denel, J. C. Fiorot, and P. Huard · 1981
Earlier work this paper cites.
Problems of Monetary Management: The UK Experience
C. A. E. Goodhart · 1984
Earlier work this paper cites.
The steepest descent gravitational method for linear programming
Soo Y. Chang and Katta G. Murty · 1989
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
Splitting Randomized Stationary Policies in Total-Reward Markov Decision Processes
Eugene A. Feinberg and Uriel G. Rothblum · 2012
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Reinforcement learning with a corrupted reward channel
Tom Everitt, Victoria Krakovna, Laurent Orseau, and Shane Legg · 2017
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in Atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
Occam’s razor is insufficient to infer the preferences of irrational agents
Soren Mindermann and Stuart Armstrong · 2018
Cited alongside, same era.
A Deep Reinforced Model for Abstractive Summarization
Romain Paulus, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Categorizing Variants of Goodhart’s Law, February 2019
David Manheim and Scott Garrabrant · 2019
Cited alongside, same era.
Observational Overfitting in Reinforcement Learning
Xingyou Song, Yiding Jiang, Stephen Tu, Yilun Du, and Behnam Neyshabur · 2019
Cited alongside, same era.
Causal Campbell-Goodhart’s Law and Reinforcement Learning:
Hal Ashton · 2021
Later among the works it cites.
The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2021
Later among the works it cites.
Reward (Mis)design for autonomous driving
W. Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt, and Peter Stone · 2022
Later among the works it cites.
Reward Gaming in Conditional Text Generation
Richard Yuanzhe Pang, Vishakh Padmakumar, Thibault Sellam, Ankur P. Parikh, and He He · 2022
Later among the works it cites.
Defining and Characterizing Reward Gaming
Joar Max Viktor Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Later among the works it cites.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Quantifying Differences in Reward Functions
Adam Gleave, Michael D. Dennis, Shane Legg, Stuart Russell, and Jan Leike · 2020
Cited alongside, same era.
Specification gaming: the flip side of AI ingenuity, 2020
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Cited alongside, same era.
Consequences of misaligned AI
Simon Zhuang and Dylan Hadfield-Menell · 2020
Cited alongside, same era.
STARC: A General Framework For Quantifying Differences Between Reward Functions, September 2023a
Joar Skalse, Lucy Farnik, Sumeet Ramesh Motwani, Erik Jenner, Adam Gleave, and Alessandro Abate
Cited in the paper.
Invariance in policy optimisation and partial identifiability in reward learning
Joar Max Viktor Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave
Cited in the paper.
Goodhart’s Law and Machine Learning: A Structural Perspective
Christopher A. Hennessy and Charles A. E. Goodhart · 2023
Closest in time.
Misspecification in inverse reinforcement learning
Joar Skalse and Alessandro Abate · 2023
Closest in time.