Fetching the paper…
Reading the bibliography…
State-action value functions (i.e., Q-values) are ubiquitous in reinforcement learning (RL), giving rise to popular algorithms such as SARSA and Q-learning.
On the theory of the brownian motion
Uhlenbeck, G. E. and Ornstein, L. S · 1930
Earlier work this paper cites.
A useful theorem for nonlinear devices having gaussian inputs
Price, R · 1958
Earlier work this paper cites.
Transformations des signaux aléatoires a travers les systemes non linéaires sans mémoire
Bonnet, G · 1964
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J · 1991
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Actor-critic algorithms, 2000
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
A theoretical and empirical analysis of expected sarsa
Van Seijen, H., Van Hasselt, H., Whiteson, S., and Wiering, M · 2009
Earlier work this paper cites.
Off-policy actor-critic
Degris, T., White, M., and Sutton, R. S · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y · 2015
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Later among the works it cites.
Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning
Gu, S., Lillicrap, T., Ghahramani, Z., Turner, R. E., Schölkopf, B., and Levine, S · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Discrete sequential prediction of continuous actions for deep RL
Metz, L., Ibarz, J., Jaitly, N., and Davidson, J · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Sample efficient actor-critic with experience replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N · 2017
Later among the works it cites.
Expected policy gradients
Ciosek, K. and Whiteson, S · 2018
Closest in time.
Trust-pcl: An off-policy trust region method for continuous control
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2018
Closest in time.