Fetching the paper…
Reading the bibliography…
How can we design agents that pursue a given objective when all feedback mechanisms are influenceable by the agent? Standard RL algorithms assume a secure reward function, and can thus perform poorly in settings where agents can tamper with the reward-generating mechanism.
Convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I. Jordan, and Satinder P. Singh · 1993
Earlier work this paper cites.
Reinforcement learning: an introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Andrew G Barto and Sridhar Mahadevan · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Preference elicitation for interface optimization
Krzysztof Gajos and Daniel S Weld · 2005
Earlier work this paper cites.
TAMER: Training an agent manually via evaluative reinforcement
W Bradley Knox and Peter Stone · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework
W. Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Learning what to value
Daniel Dewey · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Delusion, survival, and intelligent agents
Mark Ring and Laurent Orseau · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom · 2014
Earlier work this paper cites.
Approval-directed agents
Paul Christiano · 2014
Earlier work this paper cites.
Reinforcement learning with value advice
Mayank Daswani, Peter Sunehag, and Marcus Hutter · 2014
Earlier work this paper cites.
COACH: learning continuous actions from corrective advice communicated by humans
Carlos Celemin and Javier Ruiz-del Solar · 2015
Earlier work this paper cites.
Abstract approval-direction
Paul Christiano · 2015
Earlier work this paper cites.
Neural programmer-interpreters
Scott Reed and Nando De Freitas · 2015
Earlier work this paper cites.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Earlier work this paper cites.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al · 2016
Cited alongside, same era.
Concrete approval-directed agents
Paul Christiano · 2016
Cited alongside, same era.
Generalizing skills with semi-supervised reinforcement learning
Chelsea Finn, Tianhe Yu, Justin Fu, Pieter Abbeel, and Sergey Levine · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang Shane Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Later among the works it cites.
Deep TAMER: Interactive agent shaping in high-dimensional state spaces
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone · 2018
Later among the works it cites.
Deep reinforcement learning from policy-dependent human feedback
Dilip Arumugam, Jun Ki Lee, Sophie Saskin, and Michael L Littman · 2019
Later among the works it cites.
Embedded agency
Abram Demski and Scott Garrabrant · 2019
Later among the works it cites.
Reward tampering problems and solutions in reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Good and safe uses of AI oracles
Stuart Armstrong and Xavier O’Rourke · 2017
Cited alongside, same era.
Reinforcement learning with a corrupted reward channel
Tom Everitt, Victoria Krakovna, Laurent Orseau, Marcus Hutter, and Shane Legg · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Cited alongside, same era.
Tom Everitt and Marcus Hutter · 2019
Later among the works it cites.
Modeling agi safety frameworks with causal influence diagrams
Tom Everitt, Ramana Kumar, Victoria Krakovna, and Shane Legg · 2019
Later among the works it cites.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Later among the works it cites.
Language as an abstraction for hierarchical deep reinforcement learning
Yiding Jiang, Shixiang Shane Gu, Kevin P Murphy, and Chelsea Finn · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Steven Kapturowski, Georg Ostrovski, Will Dabney, John Quan, and Remi Munos · 2019
Later among the works it cites.
Misleading meta-objectives and hidden incentives for distributional shift
David Krueger, Tegan Maharaj, Shane Legg, and Jan Leike · 2019
Later among the works it cites.
Detecting spiky corruption in Markov decision processes
Jason Mancuso, Tomasz Kisielewski, David Lindner, and Alok Singh · 2019
Later among the works it cites.
Making efficient use of demonstrations to solve hard exploration problems
Tom Le Paine, Caglar Gulcehre, Bobak Shahriari, Misha Denil, Matt Hoffman, Hubert Soyer, Richard Tanburn, Steven Kapturowski, Neil Rabinowitz, Duncan Williams, et al · 2019
Later among the works it cites.
Learning human objectives by evaluating hypothetical behavior
Siddharth Reddy, Anca D Dragan, Sergey Levine, Shane Legg, and Jan Leike · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Quality of uncertainty quantification for bayesian neural network inference
Jiayu Yao, Weiwei Pan, Soumya Ghosh, and Finale Doshi-Velez · 2019
Later among the works it cites.
The incentives that shape behaviour
Ryan Carey, Eric Langlois, Tom Everitt, and Shane Legg · 2020
Closest in time.
Realab: A platform for embedded agency research
Ramana Kumar, Jonathan Uesato, Richard Ngo, Tom Everitt, Victoria Krakovna, and Shane Legg · 2020
Closest in time.
From optimizing engagement to measuring value
Smitha Milli, Luca Belli, and Moritz Hardt · 2020
Closest in time.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Closest in time.