Fetching the paper…
Reading the bibliography…
Q-learning is one of the most well-known Reinforcement Learning algorithms.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
S. Thrun and A. Schwartz · 1993
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
T. G. Dietterich · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng, S. J. Russell, et al · 2000
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
S. Legg and M. Hutter · 2007
Earlier work this paper cites.
Double q-learning
H. Hasselt · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Measuring intelligence through games
T. Schaul, J. Togelius, and J. Schmidhuber · 2011
Earlier work this paper cites.
Efficient bayes-adaptive reinforcement learning using sample-based search
A. Guez, D. Silver, and P. Dayan · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning, 2013
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Bootstrapped thompson sampling and deep exploration
I. Osband and B. Van Roy · 2015
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Cited alongside, same era.
Generalization and exploration via randomized value functions
I. Osband, B. Van Roy, and Z. Wen · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
I. Osband, J. Aslanides, and A. Cassirer · 2018
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2019
Later among the works it cites.
Using self-supervised learning can improve model robustness and uncertainty
D. Hendrycks, M. Mazeika, S. Kadavath, and D. Song · 2019
Later among the works it cites.
Deep exploration via randomized value functions
I. Osband, B. Van Roy, D. J. Russo, Z. Wen, et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas · 2016
Cited alongside, same era.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, J. Menick, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, et al · 2017
Cited alongside, same era.
Parameter space noise for exploration
M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz · 2017
Cited alongside, same era.
A tutorial on thompson sampling
D. Russo, B. Van Roy, A. Kazerouni, I. Osband, and Z. Wen · 2017
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Cited alongside, same era.
Planning to explore via self-supervised world models
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak · 2020
Later among the works it cites.
Randomized value functions via multiplicative normalizing flows
A. Touati, H. Satija, J. Romoff, J. Pineau, and P. Vincent · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Later among the works it cites.
An investigation into the effect of the learning rate on overestimation bias of connectionist q-learning
Y. Chen, L. Schomaker, and M. A. Wiering · 2021
Later among the works it cites.
First return, then explore
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2021
Later among the works it cites.