Fetching the paper…
Reading the bibliography…
In reinforcement learning (RL), function approximation errors are known to easily lead to the Q-value overestimations, thus greatly reducing policy performance.
PhD thesis, King’s College, Cambridge, 1989
C. J. C. H. Watkins, Learning from delayed rewards · 1989
Earlier work this paper cites.
S. Thrun and A. Schwartz, “Issues in using function approximation for reinforcement learning,” in Proceedings of the 1993 Connectionist Models Summer School
1993
Earlier work this paper cites.
B. Sallans and G. E. Hinton, “Reinforcement learning with factored states and actions,” Journal of Machine Learning Research
2004
Earlier work this paper cites.
H. van Hasselt, “Double q-learning,” in 23rd Advances in Neural Information Processing Systems (NeurIPS 2010)
2010
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, (IROS 2012)
2012
Earlier work this paper cites.
D. Lee, B. Defourny, and W. B. Powell, “Bias-corrected q-learning to control max-operator bias in q-learning,” in 2013 IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning (ADPRL)
2013
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114
2013
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proceedings of the 31st International Conference on Machine Learning (ICML 2014)
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on Machine Learning, (ICML 2015)
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al
2016
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in Proceedings of the 33rd International Conference on Machine Learning, (ICML 2016)
2016
Earlier work this paper cites.
H. van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in Proceedings of the 30th Conference on Artificial Intelligence (AAAI 2016)
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in 4th International Conference on Learning Representations (ICLR 2016)
2016
Earlier work this paper cites.
B. O’Donoghue, R. Munos, K. Kavukcuoglu, and V. Mnih, “Combining policy gradient and q-learning,” in 4th International Conference on Learning Representations (ICLR 2016)
2016
Cited alongside, same era.
R. Fox, A. Pakman, and N. Tishby, “Taming the noise in reinforcement learning via soft updates,” in Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence (UAI 2016)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” arXiv preprint arXiv:1606.08415
2016
Cited alongside, same era.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in Proceedings of the 35th International Conference on Machine Learning (ICML 2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos, “Distributional reinforcement learning with quantile regression,” in Proceedings of the 32nd Conference on Artificial Intelligence, (AAAI 2018)
2018
Later among the works it cites.
W. Dabney, G. Ostrovski, D. Silver, and R. Munos, “Implicit quantile networks for distributional reinforcement learning,” in Proceedings of the 35th International Conference on Machine Learning (ICML 2018)
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
M. G. Bellemare, W. Dabney, and R. Munos, “A distributional perspective on reinforcement learning,” in Proceedings of the 34th International Conference on Machine Learning, (ICML 2017)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans, “Bridging the gap between value and policy based reinforcement learning,” in 30th Advances in Neural Information Processing Systems (NeurIPS 2017)
2017
Cited alongside, same era.
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep energy-based policies,” in Proceedings of the 34th International Conference on Machine Learning, (ICML 2017)
2017
Cited alongside, same era.
MIT press, 2018
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction · 2018
Cited alongside, same era.
M. Rowland, M. Bellemare, W. Dabney, R. Munos, and Y. W. Teh, “An analysis of categorical distributional reinforcement learning,” in International Conference on Artificial Intelligence and Statistics, (AISTATS 2018)
2018
Later among the works it cites.
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, D. TB, A. Muldal, N. Heess, and T. P. Lillicrap, “Distributed distributional deterministic policy gradients,” in 6th International Conference on Learning Representations, (ICLR 2018)
2018
Later among the works it cites.
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu, “IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures,” in Proceedings of the 35th International Conference on Machine Learning (ICML 2018)
2018
Later among the works it cites.
2018
Later among the works it cites.
D. Lee and W. B. Powell, “Bias-corrected q-learning with multistate extension,” IEEE Transactions on Automatic Control
2019
Later among the works it cites.
C. Lyle, M. G. Bellemare, and P. S. Castro, “A comparative analysis of expected and distributional reinforcement learning,” in Proceedings of the 33rd Conference on Artificial Intelligence (AAAI 2019)
2019
Later among the works it cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in 3rd International Conference on Learning Representations, (ICLR 2015)
2019
Later among the works it cites.
J. Duan, S. E. Li, Y. Guan, Q. Sun, and B. Cheng, “Hierarchical reinforcement learning for self-driving decision-making without reliance on labelled driving data,” IET Intelligent Transport Systems
2020
Closest in time.
W. Dabney, Z. Kurth-Nelson, N. Uchida, C. K. Starkweather, D. Hassabis, R. Munos, and M. Botvinick, “A distributional code for value in dopamine-based reinforcement learning,” Nature
2020
Closest in time.