Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E. (2009) · 2009
Cited alongside, same era.
Gradient temporal-difference learning algorithms
Maei, H. R. (2011) · 2011
Cited alongside, same era.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M. (2012) · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T. (2013) · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S. (2014) · 2014
Cited alongside, same era.
Subgaussian concentration inequalities for geometrically ergodic Markov chains
Dedecker, J. and Gouëzel, S. (2015) · 2015
Cited alongside, same era.
On TD (0) with function approximation: Concentration bounds and a centered variant with exponential convergence
Korda, N. and La, P. (2015) · 2015
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
Liu, B., Liu, J., Ghavamzadeh, M., Mahadevan, S., and Petrik, M. (2015) · 2015
Cited alongside, same era.
OpenAI Gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Dynamics of stochastic approximation with Markov iterate-dependent noise with the stability of the iterates not ensured
Original
Karmakar, P. and Bhatnagar, S. (2016) · 2016
Cited alongside, same era.
Stochastic variance reduction methods for policy evaluation
Du, S. S., Chen, J., Li, L., Xiao, L., and Zhou, D. (2017) · 2017
Cited alongside, same era.
Finite sample analyses for TD (0) with function approximation
Dalal, G., Szörényi, B., Thoppe, G., and Mannor, S. (2018a)
Cited in the paper.