Fetching the paper…
Reading the bibliography…
We tackle the issue of finding a good policy when the number of policy updates is limited.
Maximum likelihood from incomplete data via the em algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
On the limited memory bfgs method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
P. Dayan and G. E. Hinton · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
A natural policy gradient
S. Kakade · 2001
Cited alongside, same era.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schaal · 2007
Cited alongside, same era.
Natural actor-critic
J. Peters and S. Schaal · 2008
Cited alongside, same era.
Graphical models, exponential families, and variational inference
M. J. Wainwright and M. I. Jordan · 2008
Cited alongside, same era.
Policy search for motor primitives in robotics
J. Kober and J. R. Peters · 2009
Cited alongside, same era.
On a connection between importance sampling and the likelihood ratio policy gradient
T. Jie and P. Abbeel · 2010
Cited alongside, same era.
Variational inference for policy search in changing situations
G. Neumann · 2011
Later among the works it cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller · 2013
Later among the works it cites.
Learning motor skills: from algorithms to robot experiments
J. Kober · 2014
Later among the works it cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Later among the works it cites.
Counterfactual risk minimization: Learning from logged bandit feedback
A. Swaminathan and T. Joachims · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Dudík, J. Langford, and L. Li · 2011
Cited alongside, same era.
Expectation-maximization as lower bound maximization
T. Minka
Cited in the paper.
Gradient estimation using stochastic computation graphs
J. Schulman, N. Heess, T. Weber, and P. Abbeel
Cited in the paper.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel
Cited in the paper.
Continuous deep q-learning with model-based acceleration
S. Gu, T. P. Lillicrap, I. Sutskever, and S. Levine · 2016
Closest in time.
Mastering the game of go with deep neural networks and tree search
D. Silver et al · 2016
Closest in time.