Fetching the paper…
Reading the bibliography…
Potential-based reward shaping (PBRS) is a particular category of machine learning methods which aims to improve the learning speed of a reinforcement learning agent by extracting and utilizing extra knowledge while performing a task.
Watkins, Christopher J.C.H. and Dayan, Peter, “Technical Note: Q-Learning,”
1992
Earlier work this paper cites.
Williams, Ronald J., “Simple statistical gradient-following algorithms for connectionist reinforcement learning,”
1992
Earlier work this paper cites.
Sutton, Richard S. and Barto, Andrew G., “Introduction to Reinforcement Learning,”
1998
Earlier work this paper cites.
Sutton, Richard S. and McAllester, David and Singh, Satinder and Mansour, Yishay, “Policy Gradient Methods for Reinforcement Learning with Function Approximation,” in
1999
Earlier work this paper cites.
Ng, Andrew Y. and Harada, Daishi and Russell, Stuart J., “Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping,” in
1999
Earlier work this paper cites.
Peter L. Bartlett and Jonathan Baxter, “Infinite-Horizon Policy-Gradient Estimation,”
2001
Earlier work this paper cites.
Wiewiora, Eric and Cottrell, Garrison and Elkan, Charles, “Principled Methods for Advising Reinforcement Learning Agents,” in
2003
Earlier work this paper cites.
Grześ, Marek and Kudenko, Daniel, “Multigrid Reinforcement Learning with Reward Shaping,” in
2008
Cited alongside, same era.
M. Grzes and D. Kudenko, “Plan-based reward shaping for reinforcement learning,” in
2008
Cited alongside, same era.
Taylor, Matthew E. and Stone, Peter, “Transfer Learning for Reinforcement Learning Domains: A Survey,”
2009
Cited alongside, same era.
Devlin, Sam and Kudenko, Daniel, “Dynamic Potential-based Reward Shaping,” in
2012
Cited alongside, same era.
Bellemare, Marc G. and Naddaf, Yavar and Veness, Joel and Bowling, Michael, “The Arcade Learning Environment: An Evaluation Platform for General Agents,”
2013
Cited alongside, same era.
Brys, Tim and Harutyunyan, Anna and Taylor, Matthew E. and Nowé, Ann, “Policy Transfer Using Reward Shaping,” in
2015
Later among the works it cites.
Brys, Tim and Harutyunyan, Anna and Suay, Halit Bener and Chernova, Sonia and Taylor, Matthew E. and Nowé, Ann, “Reinforcement Learning from Demonstration Through Shaping,” in
2015
Later among the works it cites.
De la Cruz, Gabriel and Du, Yunshu and Irwin, James and Taylor, Matthew, “Initial Progress in Transfer for Deep Reinforcement Learning Algorithms,” in
2016
Later among the works it cites.
Sam Devlin and Daniel Kudenko, “Plan-based reward shaping for multi-agent reinforcement learning,”
2016
Later among the works it cites.
Suay, Halit Bener and Brys, Tim and Taylor, Matthew E. and Chernova, Sonia, “Learning from Demonstration for Shaping Through Inverse Reinforcement Learning,” in
2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
Harutyunyan, Anna and Devlin, Sam and Vrancx, Peter and Nowe, Ann, “Expressing Arbitrary Reward Functions As Potential-based Advice,” in
2015
Cited alongside, same era.
Later among the works it cites.
Volodymyr Mnih and Adria Puigdomenech Badia and Mehdi Mirza and Alex Graves and Timothy Lillicrap and Tim Harley and David Silver and Koray Kavukcuoglu, “Asynchronous Methods for Deep Reinforcement Learning,” in
2016
Later among the works it cites.
Babaeizadeh, Mohammad and Frosio, Iuri and Tyree, Stephen and Clemons, Jason and Kautz, Jan, “Reinforcement Learning thorugh Asynchronous Advantage Actor-Critic on a GPU,” in
2017
Later among the works it cites.