Fetching the paper…
Reading the bibliography…
In order to solve a task using reinforcement learning, it is necessary to first formalise the goal of that task as a reward function.
On-line Q-learning using connectionist systems , volume 37
Gavin A Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart Russell · 2000
Earlier work this paper cites.
Quantifying differences in reward functions, 2020
Adam Gleave, Michael Dennis, Shane Legg, Stuart Russell, and Jan Leike · 2006
Earlier work this paper cites.
Splitting randomized stationary policies in total-reward markov decision processes
Eugene A. Feinberg and Uriel G. Rothblum · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Deep reinforcement learning from human preferences, 2017
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Reinforcement learning with a corrupted reward channel
Tom Everitt, Victoria Krakovna, Laurent Orseau, Marcus Hutter, and Shane Legg · 2017
Cited alongside, same era.
Reward learning from human preferences and demonstrations in Atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Artificial Intelligence: A Modern Approach
Stuart Russell and Peter Norvig · 2020
Simon Zhuang and Dylan Hadfield-Menell · 2021
Later among the works it cites.
Calculus on MDPs: Potential shaping as a gradient, 2022
Erik Jenner, Herke van Hoof, and Adam Gleave · 2022
Later among the works it cites.
The effects of reward misspecification: Mapping and mitigating misaligned models, 2022
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Later among the works it cites.
Dynamics-aware comparison of learned reward functions
Blake Wulfe, Logan Michael Ellis, Jean Mercat, Rowan Thomas McAllister, and Adrien Gaidon · 2022
Later among the works it cites.
Misspecification in inverse reinforcement learning, 2023
Joar Skalse and Alessandro Abate · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Invariance in policy optimisation and partial identifiability in reward learning
Joar Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave
Cited in the paper.
Defining and characterizing reward hacking
Joar Skalse, Niki Howe, Krasheninnikov Dima, and David Krueger
Cited in the paper.