Fetching the paper…
Reading the bibliography…
The objective of a reinforcement learning agent is to behave so as to maximise the sum of a suitable scalar function of state: the reward.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
A dynamic allocation index for the sequential design of experiments
Gittins, J · 1974
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
Gittins, J. C · 1979
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Simple principles of metalearning
Schmidhuber, J., Zhao, J., and Wiering, M · 1996
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
Randlöv, J. and Alström, P · 1998
Earlier work this paper cites.
Learning to learn: Introduction and overview
Thrun, S. and Pratt, L · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. J · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S., and Mansour, Y · 2000
Earlier work this paper cites.
Bayesian surprise attracts human attention
Itti, L. and Baldi, P. F · 2006
Earlier work this paper cites.
An analytic solution to discrete bayesian reinforcement learning
Poupart, P., Vlassis, N., Hoey, J., and Regan, K · 2006
Earlier work this paper cites.
Should i stay or should i go? how the human brain manages the trade-off between exploitation and exploration
Cohen, J. D., McClure, S. M., and Yu, A. J · 2007
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Cited alongside, same era.
Where do rewards come from
Singh, S., Lewis, R. L., and Barto, A. G · 2009
Cited alongside, same era.
Intrinsically motivated reinforcement learning: An evolutionary perspective
Singh, S., Lewis, R. L., Barto, A. G., and Sorg, J · 2010
Cited alongside, same era.
Reward design via online gradient ascent
Sorg, J., Lewis, R. L., and Singh, S · 2010
Cited alongside, same era.
Reinforcement active learning hierarchical loops
Gordon, G. and Ahissar, E · 2011
Cited alongside, same era.
Functions and mechanisms of intrinsic motivations
Mirolli, M. and Baldassarre, G · 2013
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Later among the works it cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2018
Later among the works it cites.
Infobot: Transfer and exploration via the information bottleneck
Goyal, A., Islam, R., Strouse, D., Ahmed, Z., Larochelle, H., Botvinick, M., Bengio, Y., and Levine, S · 2018
Later among the works it cites.
Discovery of predictive representations with a network of general value functions, 2018
Schlegel, M., Patterson, A., White, A., and White, M · 2018
Later among the works it cites.
The importance of sampling inmeta-reinforcement learning
Stadie, B., Yang, G., Houthooft, R., Chen, P., Duan, Y., Wu, Y., Abbeel, P., and Sutskever, I · 2018
Later among the works it cites.
On learning intrinsic rewards for policy gradient methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Humans use directed and random exploration to solve the explore–exploit dilemma
Wilson, R. C., Geana, A., White, J. M., Ludvig, E. A., and Cohen, J. D · 2014
Cited alongside, same era.
Expressing arbitrary reward functions as potential-based advice
Harutyunyan, A., Devlin, S., Vrancx, P., and Nowe, A · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Faulty reward functions in the wild
Clark, J. and Amodei, D · 2016
Cited alongside, same era.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Cited alongside, same era.
Deep learning for reward design to improve monte carlo tree search in atari games
Guo, X., Singh, S., Lewis, R., and Lee, H · 2016
Cited alongside, same era.
Zheng, Z., Oh, J., and Singh, S · 2018
Later among the works it cites.
Learning to understand goal specifications by modelling reward
Bahdanau, D., Hill, F., Leike, J., Hughes, E., Hosseini, A., Kohli, P., and Grefenstette, E · 2019
Closest in time.
Meta-learning via learned loss
Bechtle, S., Molchanov, A., Chebotar, Y., Grefenstette, E., Righetti, L., Sukhatme, G., and Meier, F · 2019
Closest in time.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Clavera, I., Nagabandi, A., Liu, S., Fearing, R. S., Abbeel, P., Levine, S., and Finn, C · 2019
Closest in time.
Reconciling novelty and complexity through a rational analysis of curiosity
Dubey, R. and Griffiths, T. L · 2019
Closest in time.
Improving generalization in meta reinforcement learning using learned objectives
Kirsch, L., van Steenkiste, S., and Schmidhuber, J · 2019
Closest in time.
Adapting behaviour via intrinsic reward: A survey and empirical study
Linke, C., Ady, N. M., White, M., Degris, T., and White, A · 2019
Closest in time.
Meta-learning update rules for unsupervised representation learning
Metz, L., Maheswaranathan, N., Cheung, B., and Sohl-Dickstein, J · 2019
Closest in time.
Discovery of useful questions as auxiliary tasks
Veeriah, V., Hessel, M., Xu, Z., Rajendran, J., Lewis, R. L., Oh, J., van Hasselt, H. P., Silver, D., and Singh, S · 2019
Closest in time.
Learning a prior over intent via meta-inverse reinforcement learning
Xu, K., Ratner, E., Dragan, A., Levine, S., and Finn, C · 2019
Closest in time.