Fetching the paper…
Reading the bibliography…
Learning to evaluate and improve policies is a core problem of Reinforcement Learning (RL).
Off-policy policy gradient with state distribution correction
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E. (2019) · 1904
Earlier work this paper cites.
Algaedice: Policy gradient from arbitrary experience
Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D. (2019) · 1912
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Conditional Markov processes
Stratonovich, R. (1960) · 1960
Earlier work this paper cites.
Evolutionsstrategie - Optimierung technischer Systeme nach Prinzipien der biologischen Evolution. Dissertation
Rechenberg, I. (1971) · 1973
Earlier work this paper cites.
Temporal Credit Assignment in Reinforcement Learning
Sutton, R. S. (1984) · 1984
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Reinforcement learning in Markovian and non-Markovian environments
Schmidhuber, J. (1991) · 1991
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
Schmidhuber, J. (1992) · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. (1999) · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J. (2001) · 2001
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S. (2001) · 2001
Earlier work this paper cites.
Harb, J., Schaul, T., Precup, D., and Bacon, P.-L. (2020) · 2002
Earlier work this paper cites.
Parameter-based value functions
Faccio, F., Kirsch, L., and Schmidhuber, J. (2020) · 2006
Earlier work this paper cites.
Natural actor-critic
Peters, J. and Schaal, S. (2008) · 2008
Earlier work this paper cites.
Policy gradients with parameter-based exploration for control
Sehnke, F., Osendorfer, C., Rückstieß, T., Graves, A., Peters, J., and Schmidhuber, J. (2008) · 2008
Cited alongside, same era.
Learning bounds for importance weighting
Cortes, C., Mansour, Y., and Mohri, M. (2010) · 2010
Cited alongside, same era.
Evolving neural networks in compressed weight space
Koutnik, J., Gomez, F., and Schmidhuber, J. (2010) · 2010
Cited alongside, same era.
Parameter-exploring policy gradients
Sehnke, F., Osendorfer, C., Rückstieß, T., Graves, A., Peters, J., and Schmidhuber, J. (2010) · 2010
Cited alongside, same era.
Tang, H., Meng, Z., Hao, J., Chen, C., Graves, D., Li, D., Yu, C., Mao, H., Liu, W., Yang, Y., et al. (2020) · 2010
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Later among the works it cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Later among the works it cites.
Schmidhuber, J. (2015) · 2015
Later among the works it cites.
Scalable bayesian optimization using deep neural networks
Snoek, J., Rippel, O., Swersky, K., Kiros, R., Satish, N., Sundaram, N., Patwary, M. M. A., Prabhat, P., and Adams, R. P. (2015) · 2015
Later among the works it cites.
Sample efficient actor-critic with experience replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N. (2016) · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Cited alongside, same era.
Off-policy actor-critic
Degris, T., White, M., and Sutton, R. S. (2012) · 2012
Cited alongside, same era.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P. (2012) · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J. (2013) · 2013
Cited alongside, same era.
Efficient sample reuse in policy gradients with parameter-based exploration
Zhao, T., Hachiya, H., Tangkaratt, V., Morimoto, J., and Sugiyama, M. (2013) · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I. (2017) · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y., Mansimov, E., Grosse, R. B., Liao, S., and Ba, J. (2017) · 2017
Later among the works it cites.
Spinning Up in Deep Reinforcement Learning
Achiam, J. (2018) · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Van Hoof, H., and Meger, D. (2018) · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
An off-policy policy gradient theorem using emphatic weightings
Imani, E., Graves, E., and White, M. (2018) · 2018
Later among the works it cites.
Simple random search of static linear policies is competitive for reinforcement learning
Mania, H., Guy, A., and Recht, B. (2018) · 2018
Later among the works it cites.
Policy optimization via importance sampling
Metelli, A. M., Papini, M., Faccio, F., and Restelli, M. (2018) · 2018
Later among the works it cites.
Wang, T., Zhu, J.-Y., Torralba, A., and Efros, A. A. (2018) · 2018
Later among the works it cites.