Fetching the paper…
Reading the bibliography…
We propose a new objective, the counterfactual objective, unifying existing objectives for off-policy policy gradient algorithms in the continuing reinforcement learning (RL) setting.
Off-policy policy gradient with state distribution correction
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E. (2019) · 1904
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. (1992) · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B. (1997) · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Simulation-based optimization of markov reward processes
Marbach, P. and Tsitsiklis, J. N. (2001) · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S. (2001) · 2001
Earlier work this paper cites.
A function approximation approach to estimation of policy gradient for pomdp with structured policies
Yu, H. (2005) · 2005
Earlier work this paper cites.
Non-negative matrices and Markov chains
Seneta, E. (2006) · 2006
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint
Borkar, V. S. (2009) · 2009
Earlier work this paper cites.
A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Sutton, R. S., Maei, H. R., and Szepesvári, C. (2009) · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E. (2010) · 2010
Earlier work this paper cites.
Degris, T., White, M., and Sutton, R. S. (2012) · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Policy gradient methods
Silver, D. (2015) · 2015
Cited alongside, same era.
Sample-efficient deep reinforcement learning for dialog control
Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning
Gu, S. S., Lillicrap, T., Turner, R. E., Ghahramani, Z., Schölkopf, B., and Levine, S. (2017) · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017) · 2017
Later among the works it cites.
Consistent on-line off-policy evaluation
Hallak, A. and Mannor, S. (2017) · 2017
Later among the works it cites.
Unifying task specification in reinforcement learning
White, M. (2017) · 2017
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asadi, K. and Williams, J. D. (2016) · 2016
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Generalized emphatic temporal difference learning: Bias-variance analysis
Hallak, A., Tamar, A., Munos, R., and Mannor, S. (2016) · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 2016
Cited alongside, same era.
Combining policy gradient and q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V. (2016) · 2016
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
Sutton, R. S., Mahmood, A. R., and White, M. (2016) · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N. (2016) · 2016
Cited alongside, same era.
Fujimoto, S., van Hoof, H., and Meger, D. (2018) · 2018
Later among the works it cites.
Online off-policy prediction
Ghiassian, S., Patterson, A., White, M., Sutton, R. S., and White, A. (2018) · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
An off-policy policy gradient theorem using emphatic weightings
Imani, E., Graves, E., and White, M. (2018) · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D. (2018) · 2018
Later among the works it cites.
Convergent actor-critic algorithms under off-policy training and function approximation
Maei, H. R. (2018) · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction (2nd Edition)
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
On convergence of emphatic temporal-difference learning
Yu, H. (2015) · 2018
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Gelada, C. and Bellemare, M. G. (2019) · 2019
Closest in time.