Fetching the paper…
Reading the bibliography…
Improving the sample efficiency in reinforcement learning has been a long-standing research problem.
Optimal finite-sum smooth non-convex optimization with sarah
Lam M Nguyen, Marten van Dijk, Dzung T Phan, Phuong Ha Nguyen, Tsui-Wei Weng, and Jayant R Kalagnanam · 1901
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
On measures of entropy and information
Alfréd Rényi et al · 1961
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Policy gradients with parameter-based exploration for control
Frank Sehnke, Christian Osendorfer, Thomas Rückstieß, Alex Graves, Jan Peters, and Jürgen Schmidhuber · 2008
Earlier work this paper cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2008
Earlier work this paper cites.
Learning bounds for importance weighting
Corinna Cortes, Yishay Mansour, and Mehryar Mohri · 2010
Earlier work this paper cites.
Parameter-exploring policy gradients
Frank Sehnke, Christian Osendorfer, Thomas Rückstieß, Alex Graves, Jan Peters, and Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Analysis and improvement of policy gradient estimation
Tingting Zhao, Hirotaka Hachiya, Gang Niu, and Masashi Sugiyama · 2011
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Adaptive step-size for policy gradient methods
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2013
Earlier work this paper cites.
Efficient sample reuse in policy gradients with parameter-based exploration
Tingting Zhao, Hirotaka Hachiya, Voot Tangkaratt, Jun Morimoto, and Masashi Sugiyama · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tong Zhang · 2014
Cited alongside, same era.
Stopwasting my gradients: Practical svrg
Reza Harikandeh, Mohamed Osama Ahmed, Alim Virani, Mark Schmidt, Jakub Konečnỳ, and Scott Sallinen · 2015
Cited alongside, same era.
Learning contact-rich manipulation skills with guided policy search
A simple proximal stochastic gradient method for nonsmooth nonconvex optimization
Zhize Li and Jian Li · 2018
Later among the works it cites.
Policy optimization via importance sampling
Alberto Maria Metelli, Matteo Papini, Francesco Faccio, and Marcello Restelli · 2018
Later among the works it cites.
Stochastic variance-reduced policy gradient
Matteo Papini, Damiano Binaghi, Giuseppe Canonaco, Matteo Pirotta, and Marcello Restelli · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard Turner, Zoubin Ghahramani, and Sergey Levine · 2018
Later among the works it cites.
Spiderboost: A class of faster variance-reduced algorithms for nonconvex optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Levine, Nolan Wagener, and Pieter Abbeel · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Variance reduction for faster non-convex optimization
Zeyuan Allen-Zhu and Elad Hazan · 2016
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Cited alongside, same era.
Stochastic variance reduction methods for policy evaluation
Simon S Du, Jianshu Chen, Lihong Li, Lin Xiao, and Dengyong Zhou · 2017
Cited alongside, same era.
Non-convex finite-sum optimization via scsg methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Cited alongside, same era.
Zhe Wang, Kaiyi Ji, Yi Zhou, Yingbin Liang, and Vahid Tarokh · 2018
Later among the works it cites.
Variance reduction for policy gradient with action-dependent factorized baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Later among the works it cites.
Stochastic nested variance reduced gradient descent for nonconvex optimization
Dongruo Zhou, Pan Xu, and Quanquan Gu · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Closest in time.
Neural temporal-difference learning converges to global optima
Qi Cai, Zhuoran Yang, Jason D Lee, and Zhaoran Wang · 2019
Closest in time.
A generalization theory of gradient descent for learning over-parameterized deep relu networks
Yuan Cao and Quanquan Gu · 2019
Closest in time.
Neural proximal/trust region policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Closest in time.
Hessian aided policy gradient
Zebang Shen, Alejandro Ribeiro, Hamed Hassani, Hui Qian, and Chao Mi · 2019
Closest in time.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Closest in time.
An improved convergence analysis of stochastic variance-reduced policy gradient
Pan Xu, Felicia Gao, and Quanquan Gu · 2019
Closest in time.
Policy optimization with stochastic mirror descent
Long Yang and Yu Zhang · 2019
Closest in time.
On the global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Zhuoran Yang, Yongxin Chen, Mingyi Hong, and Zhaoran Wang · 2019
Closest in time.
Policy optimization via stochastic recursive gradient algorithm, 2019
Huizhuo Yuan, Chris Junchi Li, Yuhao Tang, and Yuren Zhou · 2019
Closest in time.