Fetching the paper…
Reading the bibliography…
This paper investigates the problem of interactively learning behaviors communicated by a human teacher using positive and negative feedback.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A.G., Sutton, R.S., and Anderson, C.W · 1983
Earlier work this paper cites.
A teaching method for reinforcement learning
Clouse, Jeffery A and Utgoff, Paul E · 1992
Earlier work this paper cites.
Q-learning
Watkins, Christopher J. C. H. and Dayan, Peter · 1992
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Puterman, Martin L · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, Leemon · 1995
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, Richard S, McAllester, David A, Singh, Satinder P, Mansour, Yishay, et al · 1999
Earlier work this paper cites.
A social reinforcement learning agent
Isbell, Charles, Shelton, Christian R, Kearns, Michael, Singh, Satinder, and Stone, Peter · 2001
Earlier work this paper cites.
Giving advice about preferred actions to reinforcement learners via knowledge-based kernel regression
Maclin, Richard, Shavlik, Jude, Torrey, Lisa, Walker, Trevor, and Wild, Edward · 2005
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance
Thomaz, Andrea Lockerd and Breazeal, Cynthia · 2006
Cited alongside, same era.
Robot learning via socially guided exploration
Thomaz, Andrea L and Breazeal, Cynthia · 2007
Cited alongside, same era.
Teachable robots: Understanding human teaching behavior to build more effective robot learners
Thomaz, Andrea L and Breazeal, Cynthia · 2008
Cited alongside, same era.
A survey of robot learning from demonstration
Argall, Brenna D, Chernova, Sonia, Veloso, Manuela, and Browning, Brett · 2009
Cited alongside, same era.
Natural actor–critic algorithms
Bhatnagar, Shalabh, Sutton, Richard S, Ghavamzadeh, Mohammad, and Lee, Mark · 2009
Cited alongside, same era.
Combining manual feedback with subsequent MDP reward signals for reinforcement learning
Online human training of a myoelectric prosthesis controller via actor-critic reinforcement learning
Pilarski, Patrick M, Dawson, Michael R, Degris, Thomas, Fahimi, Farbod, Carey, Jason P, and Sutton, Richard S · 2011
Later among the works it cites.
How humans teach agents
Knox, W Bradley, Glass, Brian D, Love, Bradley C, Maddox, W Todd, and Stone, Peter · 2012
Later among the works it cites.
Learning from human-generated reward
Knox, William Bradley · 2012
Later among the works it cites.
Policy shaping: Integrating human feedback with reinforcement learning
Griffith, Shane, Subramanian, Kaushik, Scholz, Jonathan, Isbell, Charles, and Thomaz, Andrea L · 2013
Later among the works it cites.
Learning non-myopically from human-generated reward
Knox, W Bradley and Stone, Peter · 2013
Later among the works it cites.
Training a robot via human feedback: A case study
Knox, W Bradley, Stone, Peter, and Breazeal, Cynthia · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knox, W. Bradley Knox and Stone, Peter · 2010
Cited alongside, same era.
Dynamic reward shaping: training a robot by voice
Tenorio-Gonzalez, Ana C, Morales, Eduardo F, and Villaseñor-Pineda, Luis · 2010
Cited alongside, same era.
Behavior modification: Principles and procedures
Miltenberger, Raymond G · 2011
Cited alongside, same era.
Interactively shaping agents via human reinforcement: The TAMER framework
Knox, W Bradley and Stone, Peter
Cited in the paper.
Interactively shaping agents via human reinforcement: The tamer framework
Knox, W Bradley and Stone, Peter
Cited in the paper.
Later among the works it cites.
Teaching with rewards and punishments: Reinforcement or communication?
Ho, Mark K, Littman, Michael L., Cushman, Fiery, and Austerweil, Joseph L · 2015
Later among the works it cites.
Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning
Loftin, Robert, Peng, Bei, MacGlashan, James, Littman, Michael L., Taylor, Matthew E., Huang, Jeff, and Roberts, David L · 2015
Later among the works it cites.