Fetching the paper…
Reading the bibliography…
The overestimation bias is one of the major impediments to accurate off-policy learning.
Cross-validatory choice and assessment of statistical predictions
Stone, M · 1974
Earlier work this paper cites.
Non-existence of unbiased estimators of ordered parameters
Ishwaei D, B., Shabma, D., and Krishnamoorthy, K · 1985
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite markov decision processes: a review
White, D · 1988
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A · 1993
Earlier work this paper cites.
The optimizer’s curse: Skepticism and postdecision surprise in decision analysis
Smith, J. E. and Winkler, R. L · 2006
Earlier work this paper cites.
Q learning in context of approximation spaces
Patnaik, K. and Anwar, S · 2008
Earlier work this paper cites.
Double q-learning
Van Hasselt, H · 2010
Earlier work this paper cites.
Speedy q-learning
Ghavamzadeh, M., Kappen, H. J., Azar, M. G., and Munos, R · 2011
Earlier work this paper cites.
An intelligent battery controller using bias-corrected q-learning
Lee, D. and Powell, W. B · 2012
Earlier work this paper cites.
The winner’s curse: Paradoxes and anomalies of economic life
Thaler, R · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Van Hasselt, H · 2013
Earlier work this paper cites.
Deep Reinforcement Learning with Double Q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Estimating maximum expected value through gaussian approximation
D’Eramo, C., Restelli, M., and Nuara, A · 2016
Cited alongside, same era.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Cited alongside, same era.
A Distributional Perspective on Reinforcement Learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Simple random search provides a competitive approach to reinforcement learning
Mania, H., Guy, A., and Recht, B · 2018
Later among the works it cites.
Striving for simplicity in off-policy deep reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2019
Later among the works it cites.
Quantile qt-opt for risk-aware vision-based robotic grasping
Bodnar, C., Li, A., Hausman, K., Pastor, P., and Kalakrishnan, M · 2019
Later among the works it cites.
Distributional deep reinforcement learning with a mixture of gaussians
Choi, Y., Lee, K., and Oh, S · 2019
Later among the works it cites.
Interleaved q-learning with partially coupled training process
He, M. and Guo, H · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Estimating the maximum expected value in continuous reinforcement learning problems
D’Eramo, C., Nuara, A., Pirotta, M., and Restelli, M · 2017
Cited alongside, same era.
Estimating the maximum expected value through upper confidence bound of likelihood
Imagaw, T. and Kaneko, T · 2017
Cited alongside, same era.
Weighted double q-learning
Zhang, Z., Pan, Z., and Kochenderfer, M. J · 2017
Cited alongside, same era.
Distributed distributional deterministic policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., Muldal, A., Heess, N., and Lillicrap, T · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Later among the works it cites.
Mixing update q-value for deep reinforcement learning
Li, Z. and Hou, X · 2019
Later among the works it cites.
Stochastic double deep q-network
Lv, P., Wang, X., Cheng, Y., and Duan, Z · 2019
Later among the works it cites.
Distributional reinforcement learning for efficient exploration
Mavrin, B., Yao, H., Kong, L., Wu, K., and Yu, Y · 2019
Later among the works it cites.
Fully parameterized quantile function for distributional reinforcement learning
Yang, D., Zhao, L., Lin, Z., Qin, T., Bian, J., and Liu, T.-Y · 2019
Later among the works it cites.
Quota: The quantile option architecture for reinforcement learning
Zhang, S. and Yao, H · 2019
Later among the works it cites.
Maxmin q-learning: Controlling the estimation bias of q-learning
Lan, Q., Pan, Y., Fyshe, A., and White, M · 2020
Closest in time.