Fetching the paper…
Reading the bibliography…
Reward shaping (RS) is a powerful method in reinforcement learning (RL) for overcoming the problem of sparse or uninformative rewards.
Coordinating the crowd: Inducing desirable equilibria in non-cooperative systems
Mguni, D.; Jennings, J.; Macua, S. V.; Sison, E.; Ceppi, S.; and de Cote, E. M. 2019 · 1901
Earlier work this paper cites.
Reward shaping via meta-learning
Zou, H.; Ren, T.; Yan, D.; Su, H.; and Zhu, J. 2019 · 1901
Earlier work this paper cites.
Cutting Your Losses: Learning Fault-Tolerant Control and Optimal Stopping under Adverse Risk
Mguni, D. 2019 · 1902
Earlier work this paper cites.
A survey of deep reinforcement learning in video games
Shao, K.; Tang, Z.; Zhu, Y.; Li, N.; and Zhao, D. 2019 · 1912
Earlier work this paper cites.
The big match
Blackwell, D.; and Ferguson, T. S. 1968 · 1968
Earlier work this paper cites.
Tirole: Game Theory
Fudenberg, D.; and Tirole, J. 1991 · 1991
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T.; Jordan, M. I.; and Singh, S. P. 1994 · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y.; Harada, D.; and Russell, S. 1999 · 1999
Earlier work this paper cites.
Optimal stopping of Markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
Tsitsiklis, J. N.; and Van Roy, B. 1999 · 1999
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
McGovern, A.; and Barto, A. G. 2001 · 2001
Earlier work this paper cites.
PlanGAN: Model-based Planning With Sparse Rewards and Multiple Goals
Charlesworth, H.; and Montana, G. 2020 · 2006
Earlier work this paper cites.
The Impact of Non-stationarity on Generalisation in Deep Reinforcement Learning
Igl, M.; Farquhar, G.; Luketina, J.; Boehmer, W.; and Whiteson, S. 2020 · 2006
Earlier work this paper cites.
Cyclic equilibria in Markov games
Zinkevich, M.; Greenwald, A.; and Littman, M. 2006 · 2006
Earlier work this paper cites.
Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Shoham, Y.; and Leyton-Brown, K. 2008 · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
Strehl, A. L.; and Littman, M. L. 2008 · 2008
Cited alongside, same era.
On the one-dimensional optimal switching problem
Bayraktar, E.; and Egami, M. 2010 · 2010
Cited alongside, same era.
Theoretical considerations of potential-based reward shaping for multi-agent systems
Devlin, S.; and Kudenko, D. 2011 · 2011
Cited alongside, same era.
An empirical study of potential-based reward shaping and advice in complex, multi-agent systems
Devlin, S.; Kudenko, D.; and Grześ, M. 2011 · 2011
Cited alongside, same era.
Adaptive algorithms and stochastic approximations , volume 22
Benveniste, A.; Métivier, M.; and Priouret, P. 2012 · 2012
Cited alongside, same era.
Approximate dynamic programming
Bertsekas, D. P. 2012 · 2012
Cited alongside, same era.
Curiosity-driven Exploration by Self-supervised Prediction
Pathak, D.; Agrawal, P.; Efros, A. A.; and Darrell, T. 2017 · 2017
Later among the works it cites.
Proximal Policy Optimization Algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Later among the works it cites.
Exploration by random network distillation
Burda, Y.; Edwards, H.; Storkey, A.; and Klimov, O. 2018 · 2018
Later among the works it cites.
Learning Parametric Closed-Loop Policies for Markov Potential Games
Macua, S. V.; Zazo, J.; and Zazo, S. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamic potential-based reward shaping
Devlin, S. M.; and Kudenko, D. 2012 · 2012
Cited alongside, same era.
Expressing arbitrary reward functions as potential-based advice
Harutyunyan, A.; Devlin, S.; Vrancx, P.; and Nowé, A. 2015 · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C.; Levine, S.; and Abbeel, P. 2015 · 2015
Cited alongside, same era.
Playing atari games with deep reinforcement learning and human checkpoint replay
Hosu, I.-A.; and Rebedea, T. 2016 · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Houthooft, R.; Chen, X.; Duan, Y.; Schulman, J.; De Turck, F.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
Policy invariance under reward transformations for multi-objective reinforcement learning
Mannion, P.; Devlin, S.; Mason, K.; Duggan, J.; and Howley, E. 2017 · 2017
Cited alongside, same era.
Mguni, D. 2018 · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Later among the works it cites.
On Learning Intrinsic Rewards for Policy Gradient Methods
Zheng, Z.; Oh, J.; and Singh, S. 2018 · 2018
Later among the works it cites.
Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping
Hu, Y.; Wang, W.; Jia, H.; Wang, Y.; Chen, Y.; Hao, J.; Wu, F.; and Fan, C. 2020 · 2020
Later among the works it cites.
Learning Intrinsic Rewards as a Bi-Level Optimization Problem
Stadie, B.; Zhang, L.; and Ba, J. 2020 · 2020
Later among the works it cites.
Policy invariant explicit shaping: an efficient alternative to reward shaping
Behboudian, P.; Satsangi, Y.; Taylor, M. E.; Harutyunyan, A.; and Bowling, M. 2021 · 2021
Closest in time.
Learning in Nonzero-Sum Stochastic Games with Potentials
Mguni, D.; Wu, Y.; Du, Y.; Yang, Y.; Wang, Z.; Li, M.; Wen, Y.; Jennings, J.; and Wang, J. 2021 · 2021
Closest in time.
Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution Networks
Wang, J.; Xu, W.; Gu, Y.; Song, W.; and Green, T. 2021 · 2021
Closest in time.
Timing is Everything: Learning to Act Selectively with Costly Actions and Budgetary Constraints
Mguni, D.; Sootla, A.; Ziomek, J.; Slumbers, O.; Dai, Z.; Shao, K.; and Wang, J. 2022 · 2022
Closest in time.