Fetching the paper…
Reading the bibliography…
The task of building general agents that perform well over a wide range of tasks has been an important goal in reinforcement learning since its inception.
Reinforcement learning for robots using neural networks
L.-J. Lin · 1992
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Incremental multi-step q-learning
J. Peng and R. J. Williams · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
M. Riedmiller · 2005
Earlier work this paper cites.
On upper-confidence bound policies for non-stationary bandit problems
A. Garivier and E. Moulines · 2008
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
A. Garivier and E. Moulines · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
T. Salimans, J. Ho, X. Chen, S. Sidor, and I. Sutskever · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
F. P. Such, V. Madhavan, E. Conti, J. Lehman, K. O. Stanley, and J. Clune · 2017
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Cited alongside, same era.
Agent57: Outperforming the atari human benchmark
A. P. Badia, B. Piot, S. Kapturowski, P. Sprechmann, A. Vitvitskyi, Z. D. Guo, and C. Blundell · 2020
Later among the works it cites.
Do recent advancements in model-based deep reinforcement learning really improve data efficiency?, 2020
K. P. Kielak · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
I. Kostrikov, D. Yarats, and R. Fergus · 2020
Later among the works it cites.
Off-policy actor-critic with shared experience replay
S. Schmitt, M. Hessel, and K. Simonyan · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. Van Hasselt, and D. Silver · 2018
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
S. Kapturowski, G. Ostrovski, J. Quan, R. Munos, and W. Dabney · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling · 2018
Cited alongside, same era.
Learning by playing solving sparse reward tasks from scratch
M. Riedmiller, R. Hafner, T. Lampe, M. Neunert, J. Degrave, T. Wiele, V. Mnih, N. Heess, and J. T. Springenberg · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Never give up: Learning directed exploration strategies
A. P. Badia, P. Sprechmann, A. Vitvitskyi, D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, et al · 2019
Cited alongside, same era.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2019
Cited alongside, same era.
M. Schwarzer, A. Anand, R. Goel, R. D. Hjelm, A. Courville, and P. Bachman · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
A. Srinivas, M. Laskin, and P. Abbeel · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare · 2021
Later among the works it cites.
High-performance large-scale image recognition without normalization
A. Brock, S. De, S. L. Smith, and K. Simonyan · 2021
Later among the works it cites.
First return, then explore
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2021
Later among the works it cites.
Muesli: Combining improvements in policy optimization
M. Hessel, I. Danihelka, F. Viola, A. Guez, S. Schmitt, L. Sifre, T. Weber, D. Silver, and H. Van Hasselt · 2021
Later among the works it cites.
Behavior from the void: Unsupervised active pre-training
H. Liu and P. Abbeel · 2021
Later among the works it cites.
Beyond target networks: Improving deep -learning with functional regularization
A. Piche, V. Thomas, J. Marino, G. M. Marconi, C. J. Pal, and M. E. Khan · 2021
Later among the works it cites.
Return-based scaling: Yet another normalisation trick for deep rl
T. Schaul, G. Ostrovski, I. Kemaev, and D. Borsa · 2021
Later among the works it cites.
Online and offline reinforcement learning by planning with a learned model
J. Schrittwieser, T. Hubert, A. Mandhane, M. Barekatain, I. Antonoglou, and D. Silver · 2021
Later among the works it cites.
Mastering atari games with limited data
W. Ye, S. Liu, T. Kurutach, P. Abbeel, and Y. Gao · 2021
Later among the works it cites.
Fast and data efficient reinforcement learning from pixels via non-parametric value approximation
A. Long, A. Blair, and H. van Hoof · 2022
Closest in time.
The phenomenon of policy churn
T. Schaul, A. Barreto, J. Quan, and G. Ostrovski · 2022
Closest in time.