Fetching the paper…
Reading the bibliography…
How to obtain good value estimation is one of the key problems in Reinforcement Learning (RL).
Computational vision and regularization theory
T. Poggio, V. Torre, and C. Koch · 1987
Earlier work this paper cites.
Learning from delayed rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
S. Thrun and A. Schwartz · 1993
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
L. Baird · 1995
Earlier work this paper cites.
Regularization theory and neural networks architectures
F. Girosi, M. Jones, and T. Poggio · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
G. J. Gordon · 1995
Earlier work this paper cites.
Adaptive critic designs
D. V. Prokhorov and D. C. Wunsch · 1997
Earlier work this paper cites.
Actor-critic–type learning algorithms for markov decision processes
V. R. Konda and V. S. Borkar · 1999
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
On regularization algorithms in learning theory
F. Bauer, S. Pereverzev, and L. Rosasco · 2007
Earlier work this paper cites.
Double q-learning
H. Hasselt · 2010
Earlier work this paper cites.
Approximate q-learning: An introduction
D. Pandey and P. Pandey · 2010
Earlier work this paper cites.
Box2d: A 2d physics engine for games
E. Catto · 2011
Earlier work this paper cites.
Batch reinforcement learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. v. Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
O. Anschel, N. Baram, and N. Shimkin · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Policy gradient algorithms
L. Weng · 2018
Later among the works it cites.
Better exploration with optimistic actor-critic
K. Ciosek, Q. Vuong, R. Loftin, and K. Hofmann · 2019
Later among the works it cites.
Pybullet gymperium
B. Ellenberger · 2019
Later among the works it cites.
On the reduction of variance and overestimation of deep q-learning
M. Sabry and A. Khalifa · 2019
Later among the works it cites.
Revisiting the softmax bellman operator: New benefits and new perspective
Z. Song, R. Parr, and L. Carin · 2019
Later among the works it cites.
Open-source implementation for sac
D. Tianhong · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Cited alongside, same era.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Cited alongside, same era.
Weighted double q-learning
Z. Zhang, Z. Pan, and M. J. Kochenderfer · 2017
Cited alongside, same era.
Distributed distributional deterministic policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, D. Tb, A. Muldal, N. Heess, and T. Lillicrap · 2018
Cited alongside, same era.
Distributional policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, D. TB, A. Muldal, N. Heess, and T. Lillicrap · 2018
Cited alongside, same era.
Model-based value estimation for efficient model-free reinforcement learning
V. Feinberg, A. Wan, I. Stoica, M. I. Jordan, J. E. Gonzalez, and S. Levine · 2018
Cited alongside, same era.
Open-source implementation for td3
S. Fujimoto · 2018
Cited alongside, same era.
Y. Wu, G. Tucker, and O. Nachum · 2019
Later among the works it cites.
DAC: the double actor-critic architecture for learning options
S. Zhang and S. Whiteson · 2019
Later among the works it cites.
Maximum entropy-regularized multi-goal reinforcement learning
R. Zhao, X. Sun, and V. Tresp · 2019
Later among the works it cites.
Regularizing model-based planning with energy-based models
R. Boney, J. Kannala, and A. Ilin · 2020
Later among the works it cites.
How to learn a useful critic? model-based action-gradient-estimator policy optimization
P. D’Oro and W. Jaśkowski · 2020
Later among the works it cites.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
A. Kuznetsov, P. Shvechikov, A. Grishin, and D. Vetrov · 2020
Later among the works it cites.
Maxmin q-learning: Controlling the estimation bias of q-learning
Q. Lan, Y. Pan, A. Fyshe, and M. White · 2020
Later among the works it cites.
Dsac: Distributional soft actor critic for risk-sensitive reinforcement learning
X. Ma, L. Xia, Z. Zhou, J. Yang, and Q. Zhao · 2020
Later among the works it cites.
A survey of regularization strategies for deep models
R. Moradi, R. Berangi, and B. Minaei · 2020
Later among the works it cites.
Softmax deep double deterministic policy gradients
L. Pan, Q. Cai, and L. Huang · 2020
Later among the works it cites.
Opac: Opportunistic actor-critic
S. Roy, S. Bakshi, and T. Maharaj · 2020
Later among the works it cites.
Reducing estimation bias via triplet-average deep deterministic policy gradient
D. Wu, X. Dong, J. Shen, and S. C. Hoi · 2020
Later among the works it cites.