Fetching the paper…
Reading the bibliography…
We introduce Robust Exploration via Clustering-based Online Density Estimation (RECODE), a non-parametric method for novelty-based exploration that estimates visitation counts for clusters of states based on their similarity in a chosen embedding space.
Remarks on some nonparametric estimates of a density function
Rosenblatt, M · 1956
Earlier work this paper cites.
On estimation of a probability density function and mode
Parzen, E · 1962
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Markov decision processes
Puterman, M. L · 1990
Earlier work this paper cites.
Instance-based utile distinctions for reinforcement learning with hidden state
McCallum, R. A · 1995
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, D. and Sen, Ś · 2002
Earlier work this paper cites.
R-max – a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. and Tennenholtz, M · 2003
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Hutter, M · 2004
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Feature reinforcement learning: Part I. unstructured MDPs
Hutter, M. et al · 2009
Earlier work this paper cites.
Best arm identification in multi-armed bandits
Audibert, J.-Y., Bubeck, S., and Munos, R · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J · 2010
Earlier work this paper cites.
Revisiting k-means: New algorithms via bayesian nonparametrics
Kulis, B. and Jordan, M. I · 2011
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, S. and Goyal, N · 2012
Earlier work this paper cites.
Kernel-based reinforcement learning on representative states
Kveton, B. and Theocharous, G · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Q-learning for history-based reinforcement learning
Daswani, M., Sunehag, P., and Hutter, M · 2013
Cited alongside, same era.
Pac optimal exploration in continuous space markov decision processes
Pazis, J. and Parr, R · 2013
Cited alongside, same era.
Skip context tree switching
Bellemare, M., Veness, J., and Talvitie, E · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Cited alongside, same era.
Practical kernel-based reinforcement learning
Barreto, A. M., Precup, D., and Pineau, J · 2016
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Later among the works it cites.
Episodic curiosity through reachability
Savinov, N., Raichuk, A., Vincent, D., Marinier, R., Pollefeys, M., Lillicrap, T., and Gelly, S · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Later among the works it cites.
Making efficient use of demonstrations to solve hard exploration problems
Gulcehre, C., Le Paine, T., Shahriari, B., Denil, M., Hoffman, M., Soyer, H., Tanburn, R., Kapturowski, S., Rabinowitz, N., Williams, D., et al · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Conditional image generation with pixelcnn decoders
Van den Oord, A., Kalchbrenner, N., Espeholt, L., Vinyals, O., Graves, A., et al · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., van den Oord, A., and Munos, R · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
Hazan, E., Kakade, S., Singh, K., and Van Soest, A · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S · 2019
Later among the works it cites.
Bootstrap latent-predictive representations for multitask reinforcement learning
Guo, Z. D., Pires, B. A., Piot, B., Grill, J.-B., Altché, F., Munos, R., and Azar, M. G · 2020
Later among the works it cites.
Reward-free exploration for reinforcement learning
Jin, C., Krishnamurthy, A., Simchowitz, M., and Yu, T · 2020
Later among the works it cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C · 2020
Later among the works it cites.
Novelty search in representational space for sample efficient exploration
Tao, R. Y., François-Lavet, V., and Pineau, J · 2020
Later among the works it cites.
Density-based bonuses on learned representations for reward-free exploration in deep reinforcement learning
Domingues, O. D., Tallec, C., Munos, R., and Valko, M · 2021
Later among the works it cites.
Exploration-driven representation learning in reinforcement learning
Erraqabi, A., Zhao, M., Machado, M. C., Bengio, Y., Sukhbaatar, S., Denoyer, L., and Lazaric, A · 2021
Later among the works it cites.
Geometric entropic exploration
Guo, Z. D., Azar, M. G., Saade, A., Thakoor, S., Piot, B., Pires, B. A., Valko, M., Mesnard, T., Lattimore, T., and Munos, R · 2021
Later among the works it cites.
Behavior from the void: Unsupervised active pre-training
Liu, H. and Abbeel, P · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Seo, Y., Chen, L., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2021
Later among the works it cites.
Byol-explore: Exploration by bootstrapped prediction
Guo, Z. D., Thakoor, S., Pîslar, M., Pires, B. A., Altché, F., Tallec, C., Saade, A., Calandriello, D., Grill, J.-B., Tang, Y., et al · 2022
Later among the works it cites.
Kapturowski, S., Campos, V., Jiang, R., Rakićević, N., van Hasselt, H., Blundell, C., and Badia, A. P · 2022
Later among the works it cites.