Fetching the paper…
Reading the bibliography…
We present a new approach for efficient exploration which leverages a low-dimensional encoding of the environment learned with a combination of model-based and model-free objectives.
A Hitchhiker’s Guide to Statistical Comparisons of Reinforcement Learning Algorithms
Colas, C., Sigaud, O., and Oudeyer, P.-Y. (2019) · 1904
Earlier work this paper cites.
Deepmdp: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G. (2019) · 1906
Earlier work this paper cites.
A nonparametric estimate of a multivariate density function
Loftsgaarden, D. O. and Quesenberry, C. P. (1965) · 1965
Earlier work this paper cites.
Model predictive control: Theory and practice - a survey
Garcia, C. E., Prett, D. M., and Morari, M. (1989) · 1989
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J. (1990) · 1990
Earlier work this paper cites.
Efficient exploration in reinforcement learning
Thrun, S. B. (1992) · 1992
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, P. (1993) · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998) · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W. (2000) · 2000
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Chentanez, N., Barto, A. G., and Singh, S. P. (2005) · 2005
Earlier work this paper cites.
Viualizing data using t-sne
van der Maaten, L. and Hinton, G. (2008) · 2008
Earlier work this paper cites.
Information-theoretic approach to interactive learning
Still, S. (2009) · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J. (2010) · 2010
Earlier work this paper cites.
Abandoning objectives: Evolution through the search for novelty alone
Lehman, J. and Stanley, K. O. (2011) · 2011
Earlier work this paper cites.
Elements of Information Theory
Cover, T. and Thomas, J. (2012) · 2012
Earlier work this paper cites.
Intrinsically motivated model learning for a developing curious agent
Hester, T. and Stone, P. (2012) · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Still, S. and Precup, D. (2012) · 2012
Cited alongside, same era.
Changing the environment based on empowerment as intrinsic motivation
Salge, C., Glackin, C., and Polani, D. (2014) · 2014
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
García, J. and Fernández, F. (2015) · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J. (2015) · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P. (2015) · 2015
Cited alongside, same era.
Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models
Recurrent environment simulators
Chiappa, S., Racaniere, S., Wierstra, D., and Mohamed, S. (2017) · 2017
Later among the works it cites.
Where to add actions in human-in-the-loop reinforcement learning
Mandel, T., Liu, Y.-E., Brunskill, E., and Popović, Z. (2017) · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H. (2017) · 2017
Later among the works it cites.
Deep Exploration via Randomized Value Functions
Osband, I., Van Roy, B., Russo, D., and Wen, Z. (2017) · 2017
Later among the works it cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stadie, B. C., Levine, S., and Abbeel, P. (2015) · 2015
Cited alongside, same era.
Deep Reinforcement Learning with Double Q-learning
van Hasselt, H., Guez, A., and Silver, D. (2015) · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Cited alongside, same era.
Openai gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Gregor, K., Rezende, D. J., and Wierstra, D. (2016) · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P. (2016) · 2016
Cited alongside, same era.
Deep Exploration via Bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B. (2016) · 2016
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T. (2017) · 2017
Later among the works it cites.
The pycolab game engine
Stepleton, T. (2017) · 2017
Later among the works it cites.
Integrating state representation learning into deep reinforcement learning
de Bruin, T., Kober, J., Tuyls, K., and Babuška, R. (2018) · 2018
Later among the works it cites.
Combined reinforcement learning via abstract representations
François-Lavet, V., Bengio, Y., Precup, D., and Pineau, J. (2018) · 2018
Later among the works it cites.
Recurrent World Models Facilitate Policy Evolution
Ha, D. and Schmidhuber, J. (2018) · 2018
Later among the works it cites.
Learning to play with intrinsically-motivated self-aware agents
Haber, N., Mrowca, D., Fei-Fei, L., and Yamins, D. L. (2018) · 2018
Later among the works it cites.
Learning Latent Dynamics for Planning from Pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. (2018) · 2018
Later among the works it cites.
Model-Based Active Exploration
Shyam, P., Jaśkowski, W., and Gomez, F. (2018) · 2018
Later among the works it cites.
Discrete Energy on Rectifiable Sets
Borodachov, S., Hardin, D., and Saff, E. (2019) · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies
Badia, A. P., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., Tieleman, O., Arjovsky, M., Pritzel, A., Bolt, A., and Blundell, C. (2020) · 2020
Closest in time.
On bonus based exploration methods in the arcade learning environment
Taiga, A. A., Fedus, W., Machado, M. C., Courville, A., and Bellemare, M. G. (2020) · 2020
Closest in time.
Planning with expectation models
Wan, Y., Abbas, Z., White, A., White, M., and Sutton, R. S. (2020) · 2020
Closest in time.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Wang, T. and Isola, P. (2020) · 2020
Closest in time.