Fetching the paper…
Reading the bibliography…
Q-learning suffers from overestimation bias, because it approximates the maximum action value using the maximum estimated action value.
Parallel and Distributed Computation: Numerical Methods , volume 23
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
Learning from Delayed Rewards
Chris Watkins · 1989
Earlier work this paper cites.
Issues in Using Function Approximation for Reinforcement Learning
Sebastian Thrun and Anton Schwartz · 1993
Earlier work this paper cites.
Asynchronous Stochastic Approximation and Q-learning
John N Tsitsiklis · 1994
Earlier work this paper cites.
Neuro-dynamic Programming , volume 5
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Order Statistics
Herbert Aron David and Haikady Navada Nagaraja · 2004
Earlier work this paper cites.
Q-learning with Linear Function Approximation
Francisco S Melo and M Isabel Ribeiro · 2007
Earlier work this paper cites.
The Many Faces of Optimism: A Unifying Approach
István Szita and András Lőrincz · 2008
Cited alongside, same era.
Reinforcement Learning in Finite MDPs: PAC Analysis
Alexander L. Strehl, Lihong Li, and Michael L. Littman · 2009
Cited alongside, same era.
Double Q-learning
Hado van Hasselt · 2010
Cited alongside, same era.
Bias-corrected Q-learning to Control Max-operator Bias in Q-learning
Donghun Lee, Boris Defourny, and Warren B. Powell · 2013
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Deep Reinforcement Learning with Double Q-learning
Hado Hado van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Averaged-DQN: Variance Reduction and Stabilization for Deep Reinforcement Learning
Oron Anschel, Nir Baram, and Nahum Shimkin · 2017
Later among the works it cites.
Weighted Double Q-learning
Zongzhang Zhang, Zhiyuan Pan, and Mykel J. Kochenderfer · 2017
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Historical Best Q-Networks for Deep Reinforcement Learning
Wenwu Yu, Rui Wang, Ruiying Li, Jing Gao, and Xiaohui Hu · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pygame learning environment
Norman Tasfi · 2016
Cited alongside, same era.
Later among the works it cites.
MinAtar: An Atari-inspired Testbed for More Efficient Reinforcement Learning Experiments
Kenny Young and Tian Tian · 2019
Later among the works it cites.