Fetching the paper…
Reading the bibliography…
Providing Reinforcement Learning agents with expert advice can dramatically improve various aspects of learning.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Reinforcement learning for mixed open-loop and closed-loop control
Eric A Hansen, Andrew G Barto, and Shlomo Zilberstein · 1996
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
Creating advice-taking reinforcement learners
Richard Maclin and Jude W. Shavlik · 1996
Earlier work this paper cites.
Model reduction techniques for computing approximately optimal solutions for markov decision processes
Thomas Dean, Robert Givan, and Sonia Leach · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Approximate equivalence of Markov decision processes
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Principled methods for advising reinforcement learning agents
Eric Wiewiora, Garrison Cottrell, and Charles Elkan · 2003
Earlier work this paper cites.
Integrating guidance into relational reinforcement learning
Kurt Driessens and Sašo Džeroski · 2004
Earlier work this paper cites.
State abstraction discovery from irrelevant state variables
Nicholas K Jong and Peter Stone · 2005
Earlier work this paper cites.
Improving action selection in mdp’s via knowledge transfer
A.A. Sherstov and P. Stone · 2005
Earlier work this paper cites.
Methods for computing state similarity in markov decision processes
Norman Ferns, Pablo Samuel Castro, Doina Precup, and Prakash Panangaden · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance
Andrea Lockerd Thomaz and Cynthia Breazeal · 2006
Earlier work this paper cites.
Safe Exploration for Reinforcement Learning
Alexander Hans, Daniel Schneegaß, Anton Maximilian Schäfer, and Steffen Udluft · 2008
Cited alongside, same era.
Interactively shaping agents via human reinforcement: The tamer framework
W Bradley Knox and Peter Stone · 2009
Cited alongside, same era.
A new learning paradigm: Learning using privileged information
Vladimir Vapnik and Akshay Vashist · 2009
Cited alongside, same era.
Reinforcement Learning Via Practice and Critique Advice
Kshitij Judah, Saikat Roy, Alan Fern, and Thomas G Dietterich · 2010
Cited alongside, same era.
On the Theory of Learnining with Privileged Information
Dmitry Pechyony and Vladimir Vapnik · 2010
Cited alongside, same era.
Safe reinforcement learning in high-risk tasks through policy improvement
Javier Garcia and Fernando Fernandez · 2011
Cited alongside, same era.
Teaching on a budget: Agents advising agents in reinforcement learning
Lisa Torrey and Matthew Taylor · 2013
Later among the works it cites.
Active reward learning
Christian Daniel, Malte Viering, Jan Metz, Oliver Kroemer, and Jan Peters · 2014
Later among the works it cites.
Learning something from nothing: Leveraging implicit human feedback strategies
Robert Loftin, Bei Peng, James MacGlashan, Michael L Littman, Matthew E Taylor, Jie Huang, and David L Roberts · 2014
Later among the works it cites.
Domain-Independent Optimistic Initialization for Reinforcement Learning
Marlos C. Machado, Sriram Srinivasan, and Michael Bowling · 2014
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Later among the works it cites.
Goal-based action priors
David Abel, David Ellis Hershkowitz, Gabriel Barth-Maron, Stephen Brawner, Kevin O’Farrell, James MacGlashan, and Stefanie Tellex · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Augmenting reinforcement learning with human feedback
W Bradley Knox and Peter Stone · 2011
Cited alongside, same era.
Blending Autonomous Exploration and Apprenticeship Learning
Thomas J. Walsh, Daniel Hewlett, and Clayton T Morrison · 2011
Cited alongside, same era.
Dynamic potential-based reward shaping
Sam Devlin and Daniel Kudenko · 2012
Cited alongside, same era.
Safe exploration of state and action spaces in reinforcement learning
Javier Garcia and Fernando Fernandez · 2012
Cited alongside, same era.
Beyond Rewards : Learning from Richer Supervision
Pradyot K.V.N · 2012
Cited alongside, same era.
Safe Exploration in Markov Decision Processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Cited alongside, same era.
Later among the works it cites.
A Comprehensive Survey on Safe Reinforcement Learning
Javier Garcia and Fernando Fernandez · 2015
Later among the works it cites.
https://www.facebook.com/help/737806312958641
How does facebook determine what topics are trending? · 2016
Later among the works it cites.
Near optimal behavior via approximate state abstraction
David Abel, D Ellis Hershkowitz, and Michael L. Littman · 2016
Later among the works it cites.
Interactive teaching strategies for agent training
Ofra Amir, Ece Kamar, Andrey Kolobov, and Barbara Grosz · 2016
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Later among the works it cites.
Combating reinforcement learning’s sisyphean curse with intrinsic fear
Zachary C Lipton, Jianfeng Gao, Lihong Li, Jianshu Chen, and Li Deng · 2016
Later among the works it cites.
Alternating optimisation and quadrature for robust reinforcement learning
Supratik Paul, Kamil Ciosek, Michael A Osborne, and Shimon Whiteson · 2016
Later among the works it cites.
A need for speed: Adapting agent action speed to improve task learning from non-expert humans
Bei Peng, James MacGlashan, Robert Loftin, Michael L Littman, David L Roberts, and Matthew E Taylor · 2016
Later among the works it cites.
Virtual humans as centaurs: Melding real and virtual
William R Swartout · 2016
Later among the works it cites.
Theoretically-Grounded Policy Advice from Multiple Teachers in Reinforcement Learning Settings with Applications to Negative Transfer
Yusen Zhan, Haitham Bou Ammar, and Matthew E. Taylor · 2016
Later among the works it cites.