Fetching the paper…
Reading the bibliography…
Off-policy learning is key to scaling up reinforcement learning as it allows to learn about a target policy from the experience generated by a different behavior policy.
Pseudo-inverses in associative rings and semigroups
Drazin, M. P · 1958
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. et al · 1995
Earlier work this paper cites.
Neuro-dynamic programming: an overview
Bertsekas, D. P. and Tsitsiklis, J. N · 1995
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J. N., Van Roy, B., et al · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
The o.d.e. method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S. and Meyn, S. P · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D · 2000
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S · 2001
Earlier work this paper cites.
Temporal-difference networks
Sutton, R. S. and Tanner, B · 2004
Earlier work this paper cites.
On the eigenvalues of a class of saddle point matrices
Benzi, M. and Simoncini, V · 2006
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Gq ( λ \lambda ): A general gradient algorithm for temporal-difference prediction learning with eligibility traces
Maei, H. R. and Sutton, R. S · 2010
Cited alongside, same era.
Temporal difference methods for general projected equations
Bertsekas, D. P · 2011
Cited alongside, same era.
Gradient temporal-difference learning algorithms
Maei, H. R · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Cited alongside, same era.
Stabilization of stochastic iterative methods for singular and nearly singular linear systems
Q ( λ \lambda ) with off-policy corrections
Harutyunyan, A., Bellemare, M. G., Stepleton, T., and Munos, R · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M · 2016
Later among the works it cites.
Stochastic variance reduction methods for saddle-point problems
Palaniappan, B. and Bach, F · 2016
Later among the works it cites.
Stochastic forward–backward splitting for monotone inclusions
Rosasco, L., Villa, S., and Vũ, B. C · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, M. and Bertsekas, D. P · 2013
Cited alongside, same era.
Optimal primal-dual methods for a class of saddle point problems
Chen, Y., Lan, G., and Ouyang, Y · 2014
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Cited alongside, same era.
Off-policy td ( λ \lambda ) with a true online equivalence
van Hasselt, H., Mahmood, A. R., and Sutton, R. S · 2014
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
Liu, B., Liu, J., Ghavamzadeh, M., Mahadevan, S., and Petrik, M · 2015
Cited alongside, same era.
Distributed policy evaluation under multiple behavior strategies
Macua, S. V., Chen, J., Zazo, S., and Sayed, A. H · 2015
Cited alongside, same era.
Introduction to reinforcement learning with function approximation
Sutton, R. S · 2015
Cited alongside, same era.
Wang, Z., Bapst, V., Heess, N., Mnih, V., Munos, R., Kavukcuoglu, K., and de Freitas, N · 2016
Later among the works it cites.
Dalal, G., Szorenyi, B., Thoppe, G., and Mannor, S · 2017
Closest in time.
Stochastic variance reduction methods for policy evaluation
Du, S. S., Chen, J., Li, L., Xiao, L., and Zhou, D · 2017
Closest in time.
Linear stochastic approximation: Constant step-size and iterate averaging
Lakshminarayanan, C. and Szepesvári, C · 2017
Closest in time.
Multi-step off-policy learning without importance sampling ratios
Mahmood, A. R., Yu, H., and Sutton, R. S · 2017
Closest in time.
Finite sample analysis of the gtd policy evaluation algorithms in markov setting
Wang, Y., Chen, W., Liu, Y., Ma, Z.-M., and Liu, T.-Y · 2017
Closest in time.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Closest in time.