Fetching the paper…
Reading the bibliography…
Efficient exploration is an unsolved problem in Reinforcement Learning which is usually addressed by reactively rewarding the agent for fortuitously encountering novel situations.
A Mathematical Theory of Communication (Parts I and II)
Shannon, C. E · 1948
Earlier work this paper cites.
On a Measure of the Information Provided by an Experiment
Lindley, D. V · 1956
Earlier work this paper cites.
On Measures of Entropy and Information
Rényi, A · 1961
Earlier work this paper cites.
Theory of Optimal Experiments Designs
Fedorov, V · 1972
Earlier work this paper cites.
Learning From Delayed Rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Making the World Differentiable: On Using Fully Recurrent Self-Supervised Neural Networks for Dynamic Reinforcement Learning and Planning in Non-Stationary Environments
Schmidhuber, J · 1990
Earlier work this paper cites.
Reinforcement Learning Architectures for Animats
Sutton, R. S · 1991
Earlier work this paper cites.
Efficient Exploration In Reinforcement Learning
Thrun, S. B · 1992
Earlier work this paper cites.
Bayesian Experimental Design: A Review
Chaloner, K. and Verdinelli, I · 1995
Earlier work this paper cites.
Reinforcement Driven Information Acquisition in Non-Deterministic Environments
Storck, J., Hochreiter, S., and Schmidhuber, J · 1995
Earlier work this paper cites.
Reinforcement Learning: A Survey
Kaelbling, L. P., Littman, M. L., and Moore, a. W · 1996
Earlier work this paper cites.
What’s Interesting?
Schmidhuber, J · 1997
Earlier work this paper cites.
Employing EM and Pool-Based Active Learning for Text Classification
McCallum, A. K. and Nigam, K · 1998
Earlier work this paper cites.
Explorations in Efficient Reinforcement Learning
Wiering, M. A · 1999
Earlier work this paper cites.
Exploring the Predictable
Schmidhuber, J · 2002
Earlier work this paper cites.
Intrinsically Motivated Reinforcement Learning
Singh, S. P., Barto, A. G., and Chentanez, N · 2005
Cited alongside, same era.
A Theoretical Analysis of Model-Based Interval Estimation
Strehl, A. L. and Littman, M. L · 2005
Cited alongside, same era.
Pure Exploration In Multi-armed Bandits Problems
Bubeck, S., Munos, R., and Stoltz, G · 2009
Cited alongside, same era.
Optimized Expected Information Gain for Nonlinear Dynamical Systems
Busetto, A. G., Ong, C. S., and Buhmann, J. M · 2009
Cited alongside, same era.
Bayesian Surprise Attracts Human Attention
Itti, L. and Baldi, P · 2009
Cited alongside, same era.
Driven by Compression Progress: A Simple Principle Explains Essential Aspects of Subjective Beauty, Novelty, Surprise, Interestingness, Attention, Curiosity, Creativity, Art, Science, Music, Jokes
Schmidhuber, J · 2009
Human-Level Control Through Deep Reinforcement Learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Later among the works it cites.
Variational Information Maximisation For Intrinsically Motivated Reinforcement Learning
Mohamed, S. and Rezende, D. J · 2015
Later among the works it cites.
Unifying Count-Based Exploration and Intrinsic Motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Later among the works it cites.
VIME: Variational Information Maximizing Exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Later among the works it cites.
Deep Exploration via Bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Later among the works it cites.
Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Closed-form Jensen-Renyi Divergence for Mixture of Gaussians and Applications to Group-wise Shape Registration
Wang, F., Syeda-Mahmood, T., Vemuri, B. C., Beymer, D., and Rangarajan, A · 2009
Cited alongside, same era.
PILCO: A Model-Based and Data-Efficient Approach to Policy Search
Deisenroth, M. and Rasmussen, C. E · 2011
Cited alongside, same era.
Planning to Be Surprised: Optimal Bayesian Exploration in Dynamic Environments
Sun, Y., Gomez, F., and Schmidhuber, J · 2011
Cited alongside, same era.
Bayesian Inference and the Parametric Bootstrap
Efron, B · 2012
Cited alongside, same era.
Intrinsically Motivated Model Learning for Developing Curious Robots
Hester, T. and Stone, P · 2012
Cited alongside, same era.
Practical Bayesian Optimization Of Machine Learning Algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Cited alongside, same era.
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Later among the works it cites.
Curiosity-Driven Exploration by Self-Supervised Prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Later among the works it cites.
Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Closest in time.
Diversity is All You Need: Learning Skills without a Reward Function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Closest in time.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Closest in time.
Model-Ensemble Trust-Region Policy Optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Closest in time.
Computational Theories of Curiosity-Driven Learning
Oudeyer, P.-Y · 2018
Closest in time.
Large-Scale Study of Curiosity-Driven Learning
Burda, Y., Edwards, H., Pathak, D., Storkey, A., Darrell, T., and Efros, A. A · 2019
Closest in time.
Self-Supervised Exploration via Disagreement
Pathak, D., Gandhi, D., and Gupta, A · 2019
Closest in time.