Fetching the paper…
Reading the bibliography…
Reward shaping is one of the most effective methods to tackle the crucial yet challenging problem of credit assignment in Reinforcement Learning (RL).
Steps toward artificial intelligence
Minsky, M · 1961
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W · 1983
Earlier work this paper cites.
The behavior of organisms: An experimental analysis
Skinner, B. F · 1990
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Robot shaping: Developing autonomous agents through learning
Dorigo, M. and Colombetti, M · 1994
Earlier work this paper cites.
Reward functions for accelerated learning
Mataric, M. J · 1994
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Autonomous shaping: Knowledge transfer in reinforcement learning
Konidaris, G. and Barto, A · 2006
Earlier work this paper cites.
Multi-task reinforcement learning: a hierarchical bayesian approach
Wilson, A., Fern, A., Ray, S., and Tadepalli, P · 2007
Earlier work this paper cites.
Bayesian multi-task reinforcement learning
Lazaric, A. and Ghavamzadeh, M · 2010
Cited alongside, same era.
Multi-task evolutionary shaping without pre-specified representations
Snel, M. and Whiteson, S · 2010
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Cited alongside, same era.
Meta-learning with memory-augmented neural networks
Training agent for first-person shooter game with actor-critic curriculum learning
Wu, Y. and Tian, Y · 2017
Later among the works it cites.
Probabilistic model-agnostic meta-learning
Finn, C., Xu, K., and Levine, S · 2018
Later among the works it cites.
Recasting gradient-based meta-learning as hierarchical bayes
Grant, E., Finn, C., Levine, S., Darrell, T., and Griffiths, T · 2018
Later among the works it cites.
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2018
Later among the works it cites.
Openai five blog
OpenAI · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T · 2016
Cited alongside, same era.
Value iteration networks
Tamar, A., Wu, Y., Thomas, G., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al · 2016
Cited alongside, same era.
Learning to act by predicting the future
Dosovitskiy, A. and Koltun, V · 2017
Cited alongside, same era.
A developmental approach to machine learning?
Smith, L. B. and Slone, L. K · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S
Cited in the paper.
One-shot visual imitation learning via meta-learning
Finn, C., Yu, T., Zhang, T., Abbeel, P., and Levine, S
Cited in the paper.
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2018
Later among the works it cites.
Bayesian model-agnostic meta-learning
Yoon, J., Kim, T., Dia, O., Kim, S., Bengio, Y., and Ahn, S · 2018
Later among the works it cites.
Combo-action: Training agent for fps game with auxiliary tasks
Huang, S., Su, H., Zhu, J., and Chen, T · 2019
Closest in time.
Hierarchical macro strategy model for moba game ai
Wu, B., Fu, Q., Liang, J., Qu, P., Li, X., Wang, L., Liu, W., Yang, W., and Liu, Y · 2019
Closest in time.