Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel reinforcement- learning algorithm consisting in a stochastic variance-reduced version of policy gradient for solving Markov Decision Processes (MDPs).
A stochastic approximation method
Robbins, Herbert and Monro, Sutton · 1951
Earlier work this paper cites.
Simulation and the monte carlo method
Rubinstein, Reuven Y Reuven Y · 1981
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovskii, Arkadii, Yudin, David Borisovich, and Dawson, Edgar Ronald · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, Doina · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, Richard S, McAllester, David A, Singh, Satinder P, and Mansour, Yishay · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, Jonathan and Bartlett, Peter L · 2001
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Weaver, Lex and Tao, Nigel · 2001
Earlier work this paper cites.
Analysis and improvement of policy gradient estimation
Zhao, Tingting, Hachiya, Hirotaka, Niu, Gang, and Sugiyama, Masashi · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, Sham M · 2002
Earlier work this paper cites.
Large scale online learning
Bottou, Léon and LeCun, Yann · 2004
Earlier work this paper cites.
Learning bounds for importance weighting
Cortes, Corinna, Mansour, Yishay, and Mohri, Mehryar · 2010
Earlier work this paper cites.
A unifying perspective of parametric policy search methods for markov decision processes
Furmston, Thomas and Barber, David · 2012
Cited alongside, same era.
Reinforcement learning for spoken dialogue systems using off-policy natural gradient method
Jurčíček, Filip · 2012
Cited alongside, same era.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Roux, Nicolas L, Schmidt, Mark, and Bach, Francis R · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, Saeed and Lan, Guanghui · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Policy gradient in lipschitz markov decision processes
Pirotta, Matteo, Restelli, Marcello, and Bascetta, Luca · 2015
Later among the works it cites.
High-confidence off-policy evaluation
Thomas, Philip S, Theocharous, Georgios, and Ghavamzadeh, Mohammad · 2015
Later among the works it cites.
Variance reduction for faster non-convex optimization
Allen-Zhu, Zeyuan and Hazan, Elad · 2016
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Yan, Chen, Xi, Houthooft, Rein, Schulman, John, and Abbeel, Pieter · 2016
Later among the works it cites.
Mini-batch semi-stochastic gradient descent in the proximal setting
Konečnỳ, Jakub, Liu, Jie, Richtárik, Peter, and Takáč, Martin · 2016
Later among the works it cites.
Stochastic variance reduction methods for saddle-point problems
Palaniappan, Balamurugan and Bach, Francis · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Johnson, Rie and Zhang, Tong · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Yurii · 2013
Cited alongside, same era.
Monte Carlo theory, methods and examples
Owen, Art B · 2013
Cited alongside, same era.
Adaptive step-size for policy gradient methods
Pirotta, Matteo, Restelli, Marcello, and Bascetta, Luca · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P and Ba, Jimmy · 2014
Cited alongside, same era.
Stopwasting my gradients: Practical svrg
Harikandeh, Reza, Ahmed, Mohamed Osama, Virani, Alim, Schmidt, Mark, Konečnỳ, Jakub, and Sallinen, Scott · 2015
Cited alongside, same era.
Incremental majorization-minimization optimization with application to large-scale machine learning
Mairal, Julien · 2015
Cited alongside, same era.
Later among the works it cites.
Stochastic optimization with variance reduction for infinite datasets with finite sum structure
Bietti, Alberto and Mairal, Julien · 2017
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
Du, Simon S., Chen, Jianshu, Li, Lihong, Xiao, Lin, and Zhou, Dengyong · 2017
Later among the works it cites.
Adaptive batch size for safe policy gradients
Papini, Matteo, Pirotta, Matteo, and Restelli, Marcello · 2017
Later among the works it cites.
Thomas, Philip S. and Brunskill, Emma · 2017
Later among the works it cites.
Stochastic variance reduction for policy gradient estimation
Xu, Tianbing, Liu, Qiang, and Peng, Jian · 2017
Later among the works it cites.
Variance reduction for policy gradient with action-dependent factorized baselines
Wu, Cathy, Rajeswaran, Aravind, Duan, Yan, Kumar, Vikash, Bayen, Alexandre M, Kakade, Sham, Mordatch, Igor, and Abbeel, Pieter · 2018
Closest in time.