Fetching the paper…
Reading the bibliography…
This paper augments the reward received by a reinforcement learning agent with potential functions in order to help the agent learn (possibly stochastic) optimal policies.
R. J. Williams and J. Peng, “Function optimization using connectionist reinforcement learning algorithms,” Connection Science , vol. 3, no. 3, pp. 241–268, 1991
1991
Earlier work this paper cites.
S. J. Bradtke, “RL applied to linear quadratic regulation,” in Advances in Neural Information Processing Systems , 1993
1993
Earlier work this paper cites.
M. Li, T. Brys, and D. Kudenko, “Introspective reinforcement learning and learning from demonstration,” in Autonomous Agents and MultiAgent Systems , 2018, pp. 1992–1994
1994
Earlier work this paper cites.
J. Randløv and P. Alstrøm, “Learning to drive a bicycle using reinforcement learning and shaping.” in International Conference on Machine Learning , 1998
1998
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in International Conference on Machine Learning , 1999
1999
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in Neural Information Processing Systems , 2000, pp. 1057–1063
2000
Earlier work this paper cites.
V. Borkar and S. Meyn, “The ODE method for convergence of stochastic approximation and reinforcement learning,” SIAM Journal on Control and Optimization , vol. 38, no. 2, 2000
2000
Earlier work this paper cites.
E. Wiewiora, G. W. Cottrell, and C. Elkan, “Principled methods for advising reinforcement learning agents,” in International Conference on Machine Learning , 2003, pp. 792–799
2003
Earlier work this paper cites.
E. Wiewiora, “Potential-based shaping and Q-value initialization are equivalent,” Journal of Artificial Intelligence Research , pp. 205–208, 2003
2003
Earlier work this paper cites.
A. L. Thomaz and C. Breazeal, “Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance,” in AAAI , 2006, pp. 1000–1005
2006
Earlier work this paper cites.
J. Asmuth, M. L. Littman, and R. Zinkov, “Potential-based shaping in model-based RL,” in AAAI , 2008, pp. 604–609
2008
Earlier work this paper cites.
S. Bhatnagar, R. S. Sutton, M. Ghavamzadeh, and M. Lee, “Natural actor–critic algorithms,” Automatica , vol. 45, no. 11, 2009
2009
Earlier work this paper cites.
W. B. Knox and P. Stone, “Combining manual feedback with subsequent MDP reward signals for reinforcement learning,” in Autonomous Agents and Multiagent Systems , 2010, pp. 5–12
2010
Cited alongside, same era.
R. Hafner and M. Riedmiller, “Reinforcement learning in feedback control,” Machine Learning , vol. 84, pp. 137–169, 2011
2011
Cited alongside, same era.
S. M. Devlin and D. Kudenko, “Dynamic potential-based reward shaping.” in Autonomous Agents and Multiagent Systems , 2012, pp. 433–440
2012
Cited alongside, same era.
S. Levine and V. Koltun, “Guided policy search,” in International Conference on Machine Learning , 2013, pp. 1–9
2013
Cited alongside, same era.
M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming . John Wiley & Sons, 2014
2014
Cited alongside, same era.
A. Eck, L.-K. Soh, S. Devlin, and D. Kudenko, “Potential-based reward shaping for finite horizon online POMDP planning,” Autonomous Agents and Multi-Agent Systems , vol. 30, no. 3, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” in International Conference on Machine Learning , 2017
2017
Later among the works it cites.
H. Tang et al. , “# Exploration: A study of count-based exploration for deep reinforcement learning,” in Advances in Neural Information Processing Systems , 2017
2017
Later among the works it cites.
M. Grześ, “Reward shaping in episodic reinforcement learning,” in Autonomous Agents and MultiAgent Systems , 2017, pp. 565–573
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. P. Kingma and J. Ba., “Adam: A method for stochastic optimization,” arXiv:1412.6980 , 2014
2014
Cited alongside, same era.
V. Mnih et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, 2015
2015
Cited alongside, same era.
A. Harutyunyan, S. Devlin, P. Vrancx, and A. Nowé, “Expressing arbitrary reward functions as potential-based advice.” in AAAI , 2015, pp. 2652–2658
2015
Cited alongside, same era.
T. P. Lillicrap et al. , “Continuous control with deep reinforcement learning,” in International Conference on Learning and Representations , 2016
2016
Cited alongside, same era.
D. Silver et al. , “Mastering the game of Go with deep neural networks and tree search,” Nature , vol. 529, no. 7587, 2016
2016
Cited alongside, same era.
V. Mnih et al. , “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning , 2016
2016
Cited alongside, same era.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” The Journal of Machine Learning Research , vol. 17, no. 1, pp. 1334–1373, 2016
2016
Cited alongside, same era.
2017
Later among the works it cites.
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement Learning with Deep Energy-Based Policies,” in International Conference on Machine Learning , 2017, pp. 1352–1361
2017
Later among the works it cites.
M. Fazel, R. Ge, S. Kakade, and M. Mesbahi, “Global convergence of policy gradient methods for the linear quadratic regulator,” in International Conference on Machine Learning , 2018
2018
Later among the works it cites.
L. Buşoniu, T. de Bruin, D. Tolić, J. Kober, and I. Palunko, “Reinforcement learning for control: Performance, stability, and deep approximators,” Annual Reviews in Control , 2018
2018
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . MIT press, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Z. Yang, K. Zhang, M. Hong, and T. Başar, “A finite sample analysis of the actor-critic algorithm,” in IEEE Conference on Decision and Control (CDC) , 2018, pp. 2759–2764
2018
Later among the works it cites.
O. Marom and B. Rosman, “Belief reward shaping in reinforcement learning,” in AAAI , 2018, pp. 3762–3769
2018
Later among the works it cites.