Fetching the paper…
Reading the bibliography…
Off-policy sampling and experience replay are key for improving sample efficiency and scaling model-free temporal difference learning methods.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O., Chow, Y., Dai, B., and Li, L. (2019) · 1906
Earlier work this paper cites.
Iterative analysis
Varga, R. S. (1962) · 1962
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J. (1992) · 1992
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. (1995) · 1995
Earlier work this paper cites.
Neuro-dynamic programming: an overview
Bertsekas, D. P. and Tsitsiklis, J. N. (1995) · 1995
Earlier work this paper cites.
Q-learning
Watkins, C. J. C. H. and Dayan, P. (2004) · 2004
Earlier work this paper cites.
van Hasselt, H., Madjiheurem, S., Hessel, M., Silver, D., Barreto, A., and Borsa, D. (2020) · 2007
Earlier work this paper cites.
Dual representations for dynamic programming and reinforcement learning
Wang, T., Bowling, M., and Schuurmans, D. (2007) · 2007
Earlier work this paper cites.
Off-policy evaluation via the regularized lagrangian
Yang, M., Nachum, O., Dai, B., Li, L., and Schuurmans, D. (2020) · 2007
Earlier work this paper cites.
Stable dual dynamic programming
Wang, T., Bowling, M., Schuurmans, D., and Lizotte, D. J. (2008) · 2008
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E. (2009) · 2009
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013) · 2013
Earlier work this paper cites.
Matrix computations
Golub, G. H. and Van Loan, C. F. (2013) · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Cited alongside, same era.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
Sutton, R. S., Mahmood, A. R., and White, M. (2016) · 2016
Cited alongside, same era.
Consistent on-line off-policy evaluation
Hallak, A. and Mannor, S. (2017) · 2017
Cited alongside, same era.
Markov chains and mixing times
Levin, D. A. and Peres, Y. (2017) · 2017
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning and the deadly triad
van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J. (2018) · 2018
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Gelada, C. and Bellemare, M. G. (2019) · 2019
Later among the works it cites.
Multi-task deep reinforcement learning with popart
Hessel, M., Soyer, H., Espeholt, L., Czarnecki, W., Schmitt, S., and van Hasselt, H. (2019) · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
van Hasselt, H., Hessel, M., and Aslanides, J. (2019) · 2019
Later among the works it cites.
Generalized off-policy actor-critic
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-step off-policy learning without importance sampling ratios
Mahmood, A. R., Yu, H., and Sutton, R. S. (2017) · 2017
Cited alongside, same era.
Unifying task specification in reinforcement learning
White, M. (2017) · 2017
Cited alongside, same era.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K. (2018) · 2018
Cited alongside, same era.
Ghiassian, S., Patterson, A., White, M., Sutton, R. S., and White, A. (2018) · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. (2018) · 2018
Cited alongside, same era.
An off-policy policy gradient theorem using emphatic weightings
Imani, E., Graves, E., and White, M. (2018) · 2018
Cited alongside, same era.
Zhang, S., Boehmer, W., and Whiteson, S. (2019) · 2019
Later among the works it cites.
RLax: Reinforcement Learning in JAX
Budden, D., Hessel, M., Quan, J., Kapturowski, S., Baumli, K., Bhupatiraju, S., Guy, A., and King, M. (2020) · 2020
Later among the works it cites.
Haiku: Sonnet for JAX
Hennigan, T., Cai, T., Norman, T., and Babuschkin, I. (2020) · 2020
Later among the works it cites.
Optax: composable gradient transformation and optimisation, in JAX!
Hessel, M., Budden, D., Viola, F., Rosca, M., Sezener, E., and Hennigan, T. (2020) · 2020
Later among the works it cites.
Constrained markov decision processes via backward value functions
Satija, H., Amortila, P., and Pineau, J. (2020) · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M., Huang, J., and Jiang, N. (2020) · 2020
Later among the works it cites.
A self-tuning actor-critic algorithm
Zahavy, T., Xu, Z., Veeriah, V., Hessel, M., Oh, J., van Hasselt, H., Silver, D., and Singh, S. (2020) · 2020
Later among the works it cites.
Emphatic algorithms for deep reinforcement learning
Jiang, R., Zahavy, T., White, A., Xu, Z., Hessel, M., Blundell, C., and van Hasselt, H. (2021) · 2021
Closest in time.