Fetching the paper…
Reading the bibliography…
We introduce a new algorithm for reinforcement learning called Maximum aposteriori Policy Optimisation (MPO) based on coordinate ascent on a relative entropy objective.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L. Littman, and Andrew P. Moore · 1996
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1997
Earlier work this paper cites.
A convergent form of approximate policy iteration
Theodore J. Perkins and Doina Precup · 2002
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
H J Kappen · 2005
Earlier work this paper cites.
General duality between optimal control and estimation
Emanuel Todorov · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and Jan Peters · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
Evangelos Theodorou, Jonas Buchli, and Stefan Schaal · 2010
Earlier work this paper cites.
Variational inference for policy search in changing situations
Gerhard Neumann · 2011
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Earlier work this paper cites.
Variational policy search via trajectory optimization
Sergey Levine and Vladlen Koltun · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Path integral guided policy search
Yevgen Chebotar, Mrinal Kalakrishnan, Ali Yahya, Adrian Li, Stefan Schaal, and Sergey Levine · 2016
Cited alongside, same era.
Hierarchical relative entropy policy search
C. Daniel, G. Neumann, O. Kroemer, and J. Peters · 2016
Cited alongside, same era.
Efficient iterative policy optimization
Nicolas Le Roux · 2016
Later among the works it cites.
Model-free preference-based reinforcement learning
Christian Wirth, Johannes Furnkranz, and Gerhard Neumann · 2016
Later among the works it cites.
Deriving and improving cma-es with information geometric trust regions
Abbas Abdolmaleki, Bob Price, Nuno Lau, Luis Paulo Reis, Gerhard Neumann, et al · 2017
Later among the works it cites.
Hindsight experience replay, 2017
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Later among the works it cites.
Emergent complexity via multi-agent competition, 2017
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks, 2016
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Guided policy search as approximate mirror descent
William Montgomery and Sergey Levine · 2016
Cited alongside, same era.
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
Kamil Ciosek and Shimon Whiteson · 2017
Later among the works it cites.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Later among the works it cites.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E. Turner, and Sergey Levine · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Emergence of locomotion behaviours in rich environments
Nicolas Heess, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, Ali Eslami, Martin Riedmiller, et al · 2017
Later among the works it cites.
The uncertainty bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Rémi Munos, and Volodymyr Mnih · 2017
Later among the works it cites.
Convergent tree-backup and retrace with function approximation
Ahmed Touati, Pierre-Luc Bacon, Doina Precup, and Pascal Vincent · 2017
Later among the works it cites.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Rémi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2017
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller · 2018
Closest in time.