Fetching the paper…
Reading the bibliography…
Many reinforcement learning algorithms use value functions to guide the search for better policies.
On the experimental attainment of optimum conditions
Box, G. E. P. and Wilson, K. B · 1951
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Sutton, R. S · 1984
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
Spall, J. C. et al · 1992
Earlier work this paper cites.
Active learning with statistical models
Cohn, D. A., Ghahramani, Z., and Jordan, M. I · 1995
Earlier work this paper cites.
Memory-based stochastic optimization
Moore, A. W. and Schneider, J · 1996
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
No free lunch theorems for optimization
Wolpert, D. H. and Macready, W. G · 1997
Earlier work this paper cites.
Optimization using surrogate objectives on a helicopter test example
Booker, A., Dennis, J. J., Frank, P., and Serafini, D · 1998
Earlier work this paper cites.
Algorithms for sensitivity analysis of markov systems through potentials and perturbation realization
Xi-Ren Cao and Yat-Wah Wan · 1998
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L · 2001
Earlier work this paper cites.
A taxonomy of global optimization methods based on response surfaces
Jones, D. R · 2001
Earlier work this paper cites.
Evolutionary optimization of computationally expensive problems via surrogate modeling
Ong, Y. S., Nair, P. B., and Keane, A. J · 2003
Cited alongside, same era.
Relative entropy policy search
Peters, J., Mulling, K., and Altun, Y · 2010
Cited alongside, same era.
Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Cited alongside, same era.
Self-adaptive surrogate-assisted covariance matrix adaptation evolution strategy
Loshchilov, I., Schoenauer, M., and Sebag, M · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Bias in natural actor-critic algorithms
Thomas, P · 2014
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Later among the works it cites.
The predictron: End-to-end learning and planning
Silver, D., van Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D., Rabinowitz, N., Barreto, A., et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Natural evolution strategies
Wierstra, D., Schaul, T., Glasmachers, T., Sun, Y., Peters, J., and Schmidhuber, J · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2016
Cited alongside, same era.
Such, F. P., Madhavan, V., Conti, E., Lehman, J., Stanley, K. O., and Clune, J · 2017
Later among the works it cites.
TreeQN and ATreeC: Differentiable tree planning for deep reinforcement learning
Farquhar, G., Rocktäschel, T., Igl, M., and Whiteson, S · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Simple random search provides a competitive approach to reinforcement learning
Mania, H., Guy, A., and Recht, B · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
The value function polytope in reinforcement learning
Dadashi, R., Taiga, A. A., Roux, N. L., Schuurmans, D., and Bellemare, M. G · 2019
Later among the works it cites.
Likelihood ratio gradient estimation for steady-state parameters
Glynn, P. W. and Olvera-Cravioto, M · 2019
Later among the works it cites.
A comparative analysis of expected and distributional reinforcement learning
Lyle, C., Bellemare, M. G., and Castro, P. S · 2019
Later among the works it cites.