Fetching the paper…
Reading the bibliography…
Stochastic variance-reduced gradient (SVRG) is an optimization method originally designed for tackling machine learning problems with a finite sum structure.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Residual algorithms : Reinforcement learning with function approximation
Baird, L. (1995) · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Technical update: Least-squares temporal difference learning
Boyan, J. (2002) · 2002
Earlier work this paper cites.
A convergent o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Sutton, R. S., Szepesvári, C., and Maei, H. R. (2008) · 2008
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E. (2009) · 2009
Earlier work this paper cites.
Temporal difference methods for general projected equations
Bertsekas, D. P. (2011) · 2011
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Roux, N. L., Schmidt, M., and Bach, F. (2012) · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Policy evaluation with temporal differences: a survey and comparison
Dann, C., Neumann, G., and Peters, J. (2014) · 2014
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S. (2014) · 2014
Cited alongside, same era.
Stop wasting my gradients: Practical svrg
Harikandeh, R., Ahmed, M. O., Virani, A., Schmidt, M., Konečný, J., and Sallinen, S. (2015) · 2015
Finite-sample analysis of proximal gradient td algorithms
Liu, B., Liu, J., Ghavamzadeh, M., Mahadevan, S., and Petrik, M. (2015) · 2015
Later among the works it cites.
Stochastic variance reduction methods for saddle-point problems
Balamurugan, P. and Bach, F. (2016) · 2016
Later among the works it cites.
Openai gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
Du, S. S., Chen, J., Li, L., Xiao, L., and Zhou, D. (2017) · 2017
Later among the works it cites.
Less than a single pass: Stochastically controlled stochastic gradient method
Lei, L. and Jordan, M. I. (2017) · 2017
Later among the works it cites.
Stochastic variance-reduced policy gradient
Papini, M., Binaghi, D., Canonaco, G., Pirotta, M., and Restelli, M. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On td(0) with function approximation: Concentration bounds and a centered variant with exponential convergence
Korda, N. and L.A., P. (2015) · 2015
Cited alongside, same era.
Convergent tree backup and retrace with function approximation
Touati, A., Bacon, P.-L., Precup, D., and Vincent, P. (2018) · 2018
Later among the works it cites.