Fetching the paper…
Reading the bibliography…
Model-free reinforcement learning algorithms, such as Q-learning, perform poorly in the early stages of learning in noisy environments, because much effort is spent unlearning biased estimates of the state-action value function.
Competitive bidding in high-risk situations
Edward C Capen, Robert V Clapp, William M Campbell, et al · 1971
Earlier work this paper cites.
Anomalies: The winner’s curse
Richard H Thaler · 1988
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Sebastian Thrun and Anton Schwartz · 1993
Earlier work this paper cites.
When the best move isn’t optimal: Q-learning with exploration
George H John · 1994
Earlier work this paper cites.
Reinforcement learning in continuous time: Advantage updating
Leemon C Baird III · 1994
Earlier work this paper cites.
On reinforcement learning of control actions in noisy and non-Markovian domains
Mark D Pendrith and C Sammut · 1994
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Advantage updating applied to a differential game
Mance E Harmon, Leemon C Baird III, and A Harry Klopf · 1995
Earlier work this paper cites.
Estimator variance in reinforcement learning: Theoretical problems and practical solutions
Mark D Pendrith and Malcolm RK Ryan · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Richard Dearden, Nir Friedman, and Stuart Russell · 1998
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Probability theory: The logic of science
Edwin T Jaynes · 2003
Cited alongside, same era.
Rational overoptimism (and other biases)
Eric Van den Steen · 2004
Cited alongside, same era.
Learning rates for Q-learning
Eyal Even-Dar and Yishay Mansour · 2004
Cited alongside, same era.
The optimizer’s curse: Skepticism and postdecision surprise in decision analysis
James E Smith and Robert L Winkler · 2006
Cited alongside, same era.
Noisy reinforcements in reinforcement learning: some case studies based on gridworlds
Álvaro Moreno, José D Martín, Emilio Soria, Rafael Magdalena, and Marcelino Martínez · 2006
Cited alongside, same era.
Linearly-solvable Markov decision problems
Emanuel Todorov · 2006
Cited alongside, same era.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altun · 2010
Later among the works it cites.
Approximate inference and stochastic optimal control
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2010
Later among the works it cites.
Speedy Q-learning
Mohammad Ghavamzadeh, Hilbert J Kappen, Mohammad G Azar, and Rémi Munos · 2011
Later among the works it cites.
Value-difference based exploration: adaptive control between epsilon-greedy and softmax
Michel Tokic and Günther Palm · 2011
Later among the works it cites.
Trading value and information in MDPs
Jonathan Rubin, Ohad Shamir, and Naftali Tishby · 2012
Later among the works it cites.
An intelligent battery controller using bias-corrected Q-learning
Donghun Lee and Warren B Powell · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Approximate Dynamic Programming: Solving the curses of dimensionality
Warren B Powell · 2007
Cited alongside, same era.
Stochastic approximation
Vivek S Borkar · 2008
Cited alongside, same era.
A theoretical and empirical analysis of Expected Sarsa
Harm Van Seijen, Hado Van Hasselt, Shimon Whiteson, and Marco Wiering · 2009
Cited alongside, same era.
Efficient computation of optimal actions
Emanuel Todorov · 2009
Cited alongside, same era.
Algorithms for reinforcement learning
Csaba Szepesvári · 2010
Cited alongside, same era.
Double Q-learning
Hado V Hasselt · 2010
Cited alongside, same era.
Later among the works it cites.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Later among the works it cites.
Dynamic policy programming
Mohammad Gheshlaghi Azar, Vicenç Gómez, and Hilbert J Kappen · 2012
Later among the works it cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih et al · 2015
Closest in time.
Deep reinforcement learning with Double Q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Closest in time.
Increasing the action gap: New operators for reinforcement learning
Marc G Bellemare, Georg Ostrovski, Arthur Guez, Philip S Thomas, and Rémi Munos · 2016
Closest in time.