Fetching the paper…
Reading the bibliography…
Reinforcement learning has been successful in training autonomous agents to accomplish goals in complex environments.
Baicen Xiao, Bhaskar Ramasubramanian, Andrew Clark, Hannaneh Hajishirzi, Linda Bushnell, and Radha Poovendran. 2019 · 1907
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan. 1992 · 1992
Earlier work this paper cites.
An Introduction to the Bootstrap
Bradley Efron and Robert Tibshirani. 1994 · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping. In International Conference on Machine Learning
Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999 · 1999
Earlier work this paper cites.
A social reinforcement learning agent. In Proceedings of the International Conference on Autonomous Agents . 377–384
Charles Isbell, Christian R Shelton, Michael Kearns, Satinder Singh, and Peter Stone. 2001 · 2001
Earlier work this paper cites.
Dueling Network Architectures for Deep Reinforcement Learning. In International Conference on Machine Learning . 1995–2003
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas. 2016 · 2003
Earlier work this paper cites.
Principled methods for advising reinforcement learning agents. In International Conference on Machine Learning . 792–799
Eric Wiewiora, Garrison W Cottrell, and Charles Elkan. 2003 · 2003
Earlier work this paper cites.
Reinforcement learning with supervision by combining multiple learnings and expert advices. In Proceedings of the American Control Conference . 4159–4164
Hyeong Soo Chang. 2006 · 2006
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance. In AAAI . 1000–1005
Andrea Lockerd Thomaz and Cynthia Breazeal. 2006 · 2006
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework. In International Conference on Knowledge Capture . 9–16
W Bradley Knox and Peter Stone. 2009 · 2009
Earlier work this paper cites.
Combining manual feedback with subsequent MDP reward signals for reinforcement learning. In Autonomous Agents and Multiagent Systems . 5–12
W Bradley Knox and Peter Stone. 2010 · 2010
Earlier work this paper cites.
Reinforcement learning in feedback control
Roland Hafner and Martin Riedmiller. 2011 · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics . 627–635
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. 2011 · 2011
Earlier work this paper cites.
Integrating reinforcement learning with human demonstrations of varying ability. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 2 . International Foundation for Autonomous Agents and Multiagent Systems, 617–624
Matthew E Taylor, Halit Bener Suay, and Sonia Chernova. 2011 · 2011
Earlier work this paper cites.
Reinforcement learning from simultaneous human and MDP reward. In Autonomous Agents and Multiagent Systems . 475–482
W Bradley Knox and Peter Stone. 2012 · 2012
Earlier work this paper cites.
Bayesian learning for neural networks . Vol. 118
Radford M Neal. 2012 · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling. 2013 · 2013
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning. In Advances in Neural Information Processing Systems . 2625–2633
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L Isbell, and Andrea L Thomaz. 2013 · 2013
Cited alongside, same era.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman. 2014 · 2014
Cited alongside, same era.
Learning potential functions and their representations for multi-task reinforcement learning
Matthijs Snel and Shimon Whiteson. 2014 · 2014
Cited alongside, same era.
Expressing Arbitrary Reward Functions as Potential-Based Advice. In AAAI . 2652–2658
Anna Harutyunyan, Sam Devlin, Peter Vrancx, and Ann Nowé. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems . 4299–4307
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Later among the works it cites.
Simple and scalable predictive uncertainty estimation using deep ensembles. In Advances in Neural Information Processing Systems . 6402–6413
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017 · 2017
Later among the works it cites.
Interactive learning from policy-dependent human feedback. In International Conference on Machine Learning . 2285–2294
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman. 2017 · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction. In International Conference on Machine Learning
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell. 2017 · 2017
Later among the works it cites.
Improving Reinforcement Learning with Confidence-Based Demonstrations.. In IJCAI . 3027–3033
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trust region policy optimization. In International Conference on Machine Learning . 1889–1897
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Cited alongside, same era.
Dropout as Bayesian Approximation: Representing Model Uncertainty in Deep Learning. In International Conference on Machine Learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016 · 2016
Cited alongside, same era.
Continuous deep Q-learning with model-based acceleration. In International Conference on Machine Learning . 2829–2838
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine. 2016 · 2016
Cited alongside, same era.
Model-based reinforcement learning for approximate optimal regulation
Rushikesh Kamalapurkar, Patrick Walters, and Warren E Dixon. 2016 · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. 2016 · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning. In International Conference on Learning and Representations
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016 · 2016
Cited alongside, same era.
Zhaodong Wang and Matthew E Taylor. 2017 · 2017
Later among the works it cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz. 2017 · 2017
Later among the works it cites.
DQN-TAMER: Human-in-the-Loop Reinforcement Learning with Intractable Feedback
Riku Arakawa, Sosuke Kobayashi, Yuya Unno, Yuta Tsuboi, and Shin-ichi Maeda. 2018 · 2018
Later among the works it cites.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures. In International Conference on Machine Learning . 1406–1415
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Later among the works it cites.
Policy Shaping with Supervisory Attention Driven Exploration. In International Conference on Intelligent Robots and Systems . IEEE, 842–847
Taylor Kessler Faulkner, Elaine Schaertl Short, and Andrea Lockerd Thomaz. 2018 · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning. In AAAI
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver. 2018 · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Later among the works it cites.
Deep TAMER: Interactive agent shaping in high-dimensional state spaces. In AAAI
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone. 2018 · 2018
Later among the works it cites.
Deep reinforcement learning from policy-dependent human feedback
Dilip Arumugam, Jun Ki Lee, Sophie Saskin, and Michael L Littman. 2019 · 2019
Later among the works it cites.
HG-DAgger: Interactive Imitation Learning with Human Experts. In International Conference on Robotics and Automation . IEEE, 8077–8083
Michael Kelly, Chelsea Sidrane, Katherine Driggs-Campbell, and Mykel J Kochenderfer. 2019 · 2019
Later among the works it cites.
Active Attention-Modified Policy Shaping: Socially Interactive Agents Track. In Autonomous Agents and MultiAgent Systems . 728–736
Taylor Kessler Faulkner, Reymundo A Gutierrez, Elaine Schaertl Short, Guy Hoffman, and Andrea L Thomaz. 2019 · 2019
Later among the works it cites.
Leveraging Human Guidance for Deep Reinforcement Learning Tasks. In Proceedings of the International Joint Conference on Artificial Intelligence . 6339–6346
Ruohan Zhang, Faraz Torabi, Lin Guan, Dana H Ballard, and Peter Stone. 2019 · 2019
Later among the works it cites.