Fetching the paper…
Reading the bibliography…
In value-based reinforcement learning methods such as deep Q-learning, function approximation errors are known to lead to overestimated value estimates and suboptimal policies.
On the theory of the brownian motion
Uhlenbeck, G. E. and Ornstein, L. S · 1930
Earlier work this paper cites.
Dynamic Programming
Bellman, R · 1957
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A · 1993
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, D. P · 1995
Earlier work this paper cites.
Estimator variance in reinforcement learning: Theoretical problems and practical solutions
Pendrith, M. D., Ryan, M. R., et al · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Singh, S., Jaakkola, T., Littman, M. L., and Szepesvári, C · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S · 2001
Earlier work this paper cites.
On actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2003
Earlier work this paper cites.
Bias and variance approximation in value function estimates
Mannor, S., Simester, D., Sun, P., and Tsitsiklis, J. N · 2007
Earlier work this paper cites.
Biasing approximate dynamic programming with a lower discount factor
Petrik, M. and Scherrer, B · 2009
Earlier work this paper cites.
A theoretical and empirical analysis of expected sarsa
Van Seijen, H., Van Hasselt, H., Whiteson, S., and Wiering, M · 2009
Earlier work this paper cites.
Double q-learning
Van Hasselt, H · 2010
Cited alongside, same era.
Mean-variance optimization in markov decision processes
Mannor, S. and Tsitsiklis, J. N · 2011
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Bias-corrected q-learning to control max-operator bias in q-learning
Lee, D., Defourny, B., and Powell, W. B · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Later among the works it cites.
Averaged-dqn: Variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Later among the works it cites.
Openai baselines
Dhariwal, P., Hesse, C., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y · 2017
Later among the works it cites.
Deep Reinforcement Learning that Matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2016
Cited alongside, same era.
Learning to play in a day: Faster deep reinforcement learning by optimality tightening
He, F. S., Liu, Y., Schwing, A. G., and Peng, J · 2016
Cited alongside, same era.
O’Donoghue, B., Osband, I., Munos, R., and Mnih, V · 2017
Later among the works it cites.
Data-efficient deep reinforcement learning for dexterous manipulation
Popov, I., Heess, N., Lillicrap, T., Hafner, R., Barth-Maron, G., Vecerik, M., Lampe, T., Tassa, Y., Erez, T., and Riedmiller, M · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y., Mansimov, E., Grosse, R. B., Liao, S., and Ba, J · 2017
Later among the works it cites.
Distributional policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., TB, D., Muldal, A., Heess, N., and Lillicrap, T · 2018
Closest in time.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Closest in time.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., and Silver, D · 2018
Closest in time.
Smoothed action value functions for learning gaussian policies
Nachum, O., Norouzi, M., Tucker, G., and Schuurmans, D · 2018
Closest in time.