Fetching the paper…
Reading the bibliography…
Reinforcement learning is a powerful learning paradigm in which agents can learn to maximize sparse and delayed reward signals.
BOXES: An experiment in adaptive control
Donald Michie and Roger A Chambers. 1968 · 1968
Earlier work this paper cites.
Shaping as a method for accelerating reinforcement learning. In Proceedings of the 1992 IEEE International Symposium on Intelligent Control . 554–559
Vijaykumar Gullapalli and Andrew G Barto. 1992 · 1992
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding. In Advances in Neural Information Processing Systems . 1038–1044
Richard S Sutton. 1996 · 1996
Earlier work this paper cites.
Learning to Drive a Bicycle Using Reinforcement Learning and Shaping. In International Conference on Machine Learning , Vol. 98. 463–471
Jette Randløv and Preben Alstrøm. 1998 · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping. In International Conference on Machine Learning , Vol. 99. 278–287
Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999 · 1999
Earlier work this paper cites.
Potential-based shaping and Q-value initialization are equivalent
Eric Wiewiora. 2003 · 2003
Earlier work this paper cites.
Eric Wiewiora, Garrison W Cottrell, and Charles Elkan. 2003 · 2003
Earlier work this paper cites.
Heuristically Accelerated Q–Learning: a new approach to speed up Reinforcement Learning. In Brazilian Symposium on Artificial Intelligence . Springer, 245–254
Reinaldo AC Bianchi, Carlos HC Ribeiro, and Anna HR Costa. 2004 · 2004
Cited alongside, same era.
Online learning of shaping rewards in reinforcement learning
Marek Grześ and Daniel Kudenko. 2010 · 2010
Cited alongside, same era.
Dynamic potential-based reward shaping. In Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems . 433–440
Sam Michael Devlin and Daniel Kudenko. 2012 · 2012
Cited alongside, same era.
Reinforcement learning from simultaneous human and MDP reward. In Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems-Volume 1 . 475–482
W Bradley Knox and Peter Stone. 2012 · 2012
Cited alongside, same era.
Multi-objectivization of reinforcement learning problems by reward shaping. In 2014 International Joint Conference on Neural Networks . IEEE, 2315–2322
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Later among the works it cites.
Multi-objectivization and ensembles of shapings in reinforcement learning
Tim Brys, Anna Harutyunyan, Peter Vrancx, Ann Nowé, and Matthew E. Taylor. 2017 · 2017
Later among the works it cites.
Belief reward shaping in reinforcement learning. In The 32nd Conference of Association for the Advancement of Artificial Intelligence
Ofir Marom and Benjamin Rosman. 2018 · 2018
Later among the works it cites.
OpenAI Five
OpenAI. 2018 · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction (second ed.)
Richard S. Sutton and Andrew G. Barto. 2018 · 2018
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tim Brys, Anna Harutyunyan, Peter Vrancx, Matthew E. Taylor, Daniel Kudenko, and Ann Nowé. 2014 · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman. 2014 · 2014
Cited alongside, same era.
Expressing Arbitrary Reward Functions as Potential-Based Advice. In The Association for the Advancement of Artificial Intelligence . 2652–2658
Anna Harutyunyan, Sam Devlin, Peter Vrancx, and Ann Nowé. 2015 · 2015
Cited alongside, same era.
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max Jaderberg, Alexander S. Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom L. Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario Wünsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver. 2019 · 2019
Later among the works it cites.