Fetching the paper…
Reading the bibliography…
Reinforcement learning in complex environments may require supervision to prevent the agent from attempting dangerous actions.
“Dynamic Programming”
Richard Bellman · 1957
Earlier work this paper cites.
“Dynamic programming and markov processes”
Ronald Howard · 1960
Earlier work this paper cites.
“Modified Policy Iteration Algorithms for Discounted Markov Decision Problems”
Martin. Puterman and Moon Shin · 1978
Earlier work this paper cites.
“Influence diagrams”
Ronald Howard and James Matheson · 1984
Earlier work this paper cites.
“Technical Note Q-Learning”
Christopher… Watkins and Peter Dayan · 1992
Earlier work this paper cites.
“Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning”
Ronald. Williams · 1992
Earlier work this paper cites.
“On the Convergence of Stochastic Iterative Dynamic Programming Algorithms”
Tommi. Jaakkola, Michael. Jordan and Satinder. Singh · 1994
Earlier work this paper cites.
“On-line Q-learning using connectionist systems”, 1994
Gavin Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
“Generalization in Reinforcement Learning: Successful Examples Using Sparse Coarse Coding”
Richard. Sutton · 1995
Earlier work this paper cites.
“Evolutionary Algorithms for Reinforcement Learning”
David. Moriarty, Alan. Schultz and John. Grefenstette · 1999
Earlier work this paper cites.
“Convergence Results for Single-Step On-Policy Reinforcement-Learning Algorithms”
Satinder. Singh, Tommi. Jaakkola, Michael. Littman and Csaba Szepesvári · 2000
Cited alongside, same era.
“The Basic AI Drives”
Stephen. Omohundro · 2008
Cited alongside, same era.
“Uncertainty handling CMA-ES for reinforcement learning”
Verena Heidrich-Meisner and Christian Igel · 2009
Cited alongside, same era.
“Causality: Models, Reasoning and Inference”
Judea Pearl · 2009
Cited alongside, same era.
“Superintelligence: Paths, Dangers, Strategies”
Nick Bostrom · 2014
Cited alongside, same era.
“Fixed Point Quantization of Deep Convolutional Networks”
Darryl Lin, Sachin. Talathi and V. Annapureddy · 2016
Cited alongside, same era.
“Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates”
Shixiang Gu, Ethan Holly, Timothy. Lillicrap and Sergey Levine · 2017
Later among the works it cites.
Jan Leike et al · 2017
Later among the works it cites.
“Evolution Strategies as a Scalable Alternative to Reinforcement Learning”
Tim Salimans, Jonathan Ho, Xi Chen and Ilya Sutskever · 2017
Later among the works it cites.
“Safe Exploration in Continuous Action Spaces”
Gal Dalal et al · 2018
Later among the works it cites.
“Reinforcement Learning: An Introduction”
Richard. Sutton and Andrew. Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Laurent Orseau and Stuart Armstrong · 2016
Cited alongside, same era.
“Agent-Agnostic Human-in-the-Loop Reinforcement Learning”
David Abel, John Salvatier, Andreas Stuhlmüller and Owain Evans · 2017
Cited alongside, same era.
“Safe Model-based Reinforcement Learning with Stability Guarantees”
Felix Berkenkamp, Matteo Turchetta, Angela. Schoellig and Andreas Krause · 2017
Cited alongside, same era.
Tom Everitt, Pedro. Ortega, Elizabeth Barnes and Shane Legg · 2019
Later among the works it cites.
“Quantized Reinforcement Learning (QUARL)”
Srivatsan Krishnan et al · 2019
Later among the works it cites.
“Agent Incentives: A Causal Approach”
Tom Everitt et al · 2021
Closest in time.
“Trial without Error: Towards Safe Reinforcement Learning via Human Intervention”
William Saunders, Girish Sastry, Andreas Stuhlmüller and Owain Evans · 2069
Closest in time.