Fetching the paper…
Reading the bibliography…
Atari games have been a long-standing benchmark in the reinforcement learning (RL) community for the past decade.
Markov decision processes
Puterman, M. L · 1990
Earlier work this paper cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H. P., and Silver, D · 1990
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J · 1991
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V · 2007
Earlier work this paper cites.
On upper-confidence bound policies for non-stationary bandit problems, 2008
Garivier, A. and Moulines, E · 2008
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
Barto, A. G · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D · 2017
Earlier work this paper cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al · 2017
Earlier work this paper cites.
Ex2: Exploration with exemplar models for deep reinforcement learning
Fu, J., Co-Reyes, J., and Levine, S · 2017
Earlier work this paper cites.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., et al · 2017
Cited alongside, same era.
Sparse attentive backtracking: Long-range credit assignment in recurrent networks
Ke, N. R., Goyal, A., Bilaniuk, O., Binas, J., Charlin, L., Pal, C., and Bengio, Y · 2017
Cited alongside, same era.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., van den Oord, A., and Munos, R · 2017
Cited alongside, same era.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M · 2017
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
Observe and look further: Achieving consistent performance on atari
Pohlen, T., Piot, B., Hester, T., Azar, M. G., Horgan, D., Budden, D., Barth-Maron, G., Van Hasselt, H., Quan, J., Večerík, M., et al · 2018
Later among the works it cites.
Episodic curiosity through reachability
Savinov, N., Raichuk, A., Marinier, R., Vincent, D., Pollefeys, M., Lillicrap, T., and Gelly, S · 2018
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., Brandstetter, J., and Hochreiter, S · 2019
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2019
Later among the works it cites.
Credit assignment as a proxy for transfer in reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, O. X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Cited alongside, same era.
Variational option discovery algorithms
Achiam, J., Edwards, H., Amodei, D., and Abbeel, P · 2018
Cited alongside, same era.
Learning dexterous in-hand manipulation
Andrychowicz, M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al · 2018
Cited alongside, same era.
Playing hard exploration games by watching youtube
Aytar, Y., Pfaff, T., Budden, D., Paine, T., Wang, Z., and de Freitas, N · 2018
Cited alongside, same era.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Cited alongside, same era.
Contingency-aware exploration in reinforcement learning
Choi, J., Guo, Y., Moczulski, M., Oh, J., Wu, N., Norouzi, M., and Lee, H · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Cited alongside, same era.
Ferret, J., Marinier, R., Geist, M., and Pietquin, O · 2019
Later among the works it cites.
Generalization of reinforcement learners with working and episodic memory
Fortunato, M., Tan, M., Faulkner, R., Hansen, S., Badia, A. P., Buttimore, G., Deck, C., Leibo, J. Z., and Blundell, C · 2019
Later among the works it cites.
Hindsight credit assignment
Harutyunyan, A., Dabney, W., Mesnard, T., Azar, M. G., Piot, B., Heess, N., van Hasselt, H. P., Wayne, G., Singh, S., Precup, D., et al · 2019
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
Hung, C.-C., Lillicrap, T., Abramson, J., Wu, Y., Mirza, M., Carnevale, F., Ahuja, A., and Wayne, G · 2019
Later among the works it cites.
Sequence modeling of temporal credit assignment for episodic reinforcement learning
Liu, Y., Luo, Y., Zhong, Y., Chen, X., Liu, Q., and Peng, J · 2019
Later among the works it cites.
Adapting behaviour for learning progress, 2019
Schaul, T., Borsa, D., Ding, D., Szepesvari, D., Ostrovski, G., Dabney, W., and Osindero, S · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2019
Later among the works it cites.
Is deep reinforcement learning really superhuman on atari?
Toromanoff, M., Wirbel, E., and Moutarde, F · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Credit assignment techniques in stochastic computation graphs
Weber, T., Heess, N., Buesing, L., and Silver, D · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies
Puigdomènech Badia, A., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., Tieleman, O., Arjovsky, M., Pritzel, A., Bolt, A., and Blundell, C · 2020
Closest in time.