Fetching the paper…
Reading the bibliography…
Recent work on exploration in reinforcement learning (RL) has led to a series of increasingly complex solutions to the problem.
Discovering options for exploration by minimizing cover time
Jinnai, Y., Park, J. W., Abel, D., and Konidaris, G · 1903
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W · 1983
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Lévy flight search patterns of wandering albatrosses
Viswanathan, G. M., Afanasyev, V., Buldyrev, S., Murphy, E., Prince, P., and Stanley, H. E · 1996
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, R. and Russell, S. J · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Optimizing the success of random searches
Viswanathan, G. M., Buldyrev, S. V., Havlin, S., Da Luz, M., Raposo, E., and Stanley, H. E · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Singh, S., Jaakkola, T., Littman, M. L., and Szepesvári, C · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
McGovern, A. and Barto, A. G · 2001
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Speeding-up reinforcement learning with multi-step actions
Schoknecht, R. and Riedmiller, M · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
Stolle, M. and Precup, D · 2002
Earlier work this paper cites.
Learning rates for q-learning
Even-Dar, E. and Mansour, Y · 2003
Earlier work this paper cites.
Reinforcement learning on explicitly specified time scales
Schoknecht, R. and Riedmiller, M · 2003
Earlier work this paper cites.
Using relative novelty to identify useful temporal abstractions in reinforcement learning
Şimşek, Ö. and Barto, A. G · 2004
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
A bayesian sampling approach to exploration in reinforcement learning
Asmuth, J., Li, L., Littman, M. L., Nouri, A., and Wingate, D · 2009
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
Kolter, J. Z. and Ng, A. Y · 2009
Earlier work this paper cites.
A theoretical and empirical analysis of expected sarsa
Van Seijen, H., Van Hasselt, H., Whiteson, S., and Wiering, M · 2009
Earlier work this paper cites.
Environmental context explains lévy and brownian movement patterns of marine predators
Humphries, N. E., Queiroz, N., Dyer, J. R., Pade, N. G., Musyl, M. K., Schaefer, K. M., Fuller, D. W., Brunnschweiler, J. M., Doyle, T. K., Houghton, J. D., et al · 2010
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mulling, K., and Altun, Y · 2010
Earlier work this paper cites.
Value function approximation in reinforcement learning using the fourier basis
Konidaris, G., Osentoski, S., and Thomas, P · 2011
Cited alongside, same era.
Lévy flight and brownian search patterns of a free-ranging predator reflect different prey field characteristics
Sims, D. W., Humphries, N. E., Bradford, R. W., and Bruce, B. D · 2012
Cited alongside, same era.
Lévy flights in human behavior and cognition
Baronchelli, A. and Radicchi, F · 2013
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Hierarchical solution of markov decision processes using macro-actions
Hauskrecht, M., Meuleau, N., Kaelbling, L. P., Dean, T. L., and Boutilier, C · 2013
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, O. X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Later among the works it cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Cited alongside, same era.
Temporal abstraction in monte carlo tree search
Vafadost, M · 2013
Cited alongside, same era.
Hierarchical learning in stochastic domains: Preliminary results
Kaelbling, L. P · 2014
Cited alongside, same era.
Frame skip is a powerful parameter for learning to play atari
Braylan, A., Hollenbeck, M., Meyerson, E., and Miikkulainen, R · 2015
Cited alongside, same era.
Bayesian reinforcement learning: A survey
Ghavamzadeh, M., Mannor, S., Pineau, J., Tamar, A., et al · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Van Hasselt, H., Lanctot, M., and De Freitas, N · 2015
Cited alongside, same era.
Later among the works it cites.
Data center cooling using model-predictive control
Lazic, N., Boutilier, C., Lu, T., Wong, E., Roy, B., Ryu, M., and Imwalle, G · 2018
Later among the works it cites.
Liu, Y. and Brunskill, E · 2018
Later among the works it cites.
Eigenoption discovery through the deep successor representation
Machado, M. C., Rosenbaum, C., Guo, X., Liu, M., Tesauro, G., and Campbell, M · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A · 2018
Later among the works it cites.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M · 2018
Later among the works it cites.
Current status and future directions of lévy walk research
Reynolds, A. M · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Later among the works it cites.
Learning action representations for reinforcement learning
Chandak, Y., Theocharous, G., Kostas, J., Jordan, S., and Thomas, P · 2019
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2019
Later among the works it cites.
The termination critic
Harutyunyan, A., Dabney, W., Borsa, D., Heess, N., Munos, R., and Precup, D · 2019
Later among the works it cites.
Successor uncertainties: exploration and uncertainty in temporal difference learning
Janz, D., Hron, J., Mazur, P., Hofmann, K., Hernández-Lobato, J. M., and Tschiatschek, S · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2019
Later among the works it cites.
Transforming cooling optimization for green data center via deep reinforcement learning
Li, Y., Wen, Y., Tao, D., and Guan, K · 2019
Later among the works it cites.
The natural language of actions
Tennenholtz, G. and Mannor, S · 2019
Later among the works it cites.
On bonus based exploration methods in the arcade learning environment
Ali Taïga, A., Fedus, W., Machado, M. C., Courville, A., and Bellemare, M. G · 2020
Closest in time.
Never give up: Learning directed exploration strategies
Badia, A. P., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., Tieleman, O., Arjovsky, M., Pritzel, A., Bolt, A., and Blundell, C · 2020
Closest in time.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Warde-Farley, D., de Wiele, T. V., and Mnih, V · 2020
Closest in time.
Exploration in reinforcement learning with deep covering options
Jinnai, Y., Park, J. W., Machado, M. C., and Konidaris, G · 2020
Closest in time.