Fetching the paper…
Reading the bibliography…
We consider an agent's uncertainty about its environment and the problem of generalizing this uncertainty across observations.
Elements of information theory
Cover, T. M. and Thomas, J. A. (1991) · 1991
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J. (1991) · 1991
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Singh, S., Barto, A. G., and Chentanez, N. (2004) · 2004
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P., Kaplan, F., and Hafner, V. (2007) · 2007
Earlier work this paper cites.
Driven by compression progress
Schmidhuber, J. (2008) · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
Strehl, A. L. and Littman, M. L. (2008) · 2008
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Wainwright, M. J. and Jordan, M. I. (2008) · 2008
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
Kolter, Z. J. and Ng, A. Y. (2009) · 2009
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Lopes, M., Lang, T., Toussaint, M., and Oudeyer, P.-Y. (2012) · 2012
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
Barto, A. G. (2013) · 2013
Cited alongside, same era.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Cited alongside, same era.
Sparse adaptive dirichlet-multinomial-like processes
Hutter, M. (2013) · 2013
Cited alongside, same era.
Universal knowledge-seeking agents for stochastic environments
Orseau, L., Lattimore, T., and Hutter, M. (2013) · 2013
Cited alongside, same era.
Skip context tree switching
Bellemare, M., Veness, J., and Talvitie, E. (2014) · 2014
Cited alongside, same era.
Domain-independent optimistic initialization for reinforcement learning
Machado, M. C., Srinivasan, S., and Bowling, M. (2015) · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P. (2015) · 2015
Later among the works it cites.
Compress and control
Veness, J., Bellemare, M. G., Hutter, M., Chua, A., and Desjardins, G. (2015) · 2015
Later among the works it cites.
Increasing the action gap: New operators for reinforcement learning
Bellemare, M. G., Ostrovski, G., Guez, A., Thomas, P. S., and Munos, R. (2016) · 2016
Closest in time.
Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P. (2016) · 2016
Closest in time.
Thompson sampling is asymptotically optimal in general environments
Leike, J., Lattimore, T., Orseau, L., and Hutter, M. (2016) · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J. (2015) · 2015
Cited alongside, same era.
Laplace’s rule of succession in information geometry
Ollivier, Y. (2015) · 2015
Cited alongside, same era.
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 2016
Closest in time.
Efficient PAC-optimal exploration in concurrent, continuous state MDPs with delayed updates
Pazis, J. and Parr, R. (2016) · 2016
Closest in time.
Pixel recurrent neural networks
Van den Oord, A., Kalchbrenner, N., and Kavukcuoglu, K. (2016) · 2016
Closest in time.
Deep reinforcement learning with double Q-learning
van Hasselt, H., Guez, A., and Silver, D. (2016) · 2016
Closest in time.