Fetching the paper…
Reading the bibliography…
In many environments only a tiny subset of all states yield high reward.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, and Jan Peters · 1935
Earlier work this paper cites.
Optimal control and estimation
Robert F Stengel · 1986
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1989
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W Moore and Christopher G Atkeson · 1993
Earlier work this paper cites.
Population-based incremental learning: A method for integrating genetic search based function optimization and competitive learning
Shumeet Baluja · 1994
Earlier work this paper cites.
The” wake-sleep” algorithm for unsupervised neural networks
Geoffrey E Hinton, Peter Dayan, Brendan J Frey, and Radford M Neal · 1995
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1997
Earlier work this paper cites.
Actor-critic Algorithms
Vijaymohan Konda · 2002
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Sample-based learning and search with permanent and transient memories
David Silver, Richard S. Sutton, and Martin Müller · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
In Proceedings of the fourteenth international conference on artificial intelligence and statistics , 2011
A reduction of imitation learning and structured prediction to no-regret online learning · 2011
Earlier work this paper cites.
Variational inference for policy search in changing situations
Gerhard Neumann et al · 2011
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Cited alongside, same era.
Variational policy search via trajectory optimization
Sergey Levine and Vladlen Koltun · 2013
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Greg Wayne, David Silver, Timothy P. Lillicrap, Yuval Tassa, and Tom Erez · 2015
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Later among the works it cites.
Recurrent environment simulators
Silvia Chiappa, Sébastien Racaniere, Daan Wierstra, and Shakir Mohamed · 2017
Later among the works it cites.
Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning
Shixiang Gu, Tim Lillicrap, Richard E Turner, Zoubin Ghahramani, Bernhard Schölkopf, and Sergey Levine · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
End to end learning for self-driving cars
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al · 2016
Cited alongside, same era.
The CMA evolution strategy: A tutorial
Nikolaus Hansen · 2016
Cited alongside, same era.
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Automatic goal generation for reinforcement learning agents
David Held, Xinyang Geng, Carlos Florensa, and Pieter Abbeel · 2017
Later among the works it cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S Fearing, and Sergey Levine · 2017
Later among the works it cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aaron van den Oord, and Rémi Munos · 2017
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Ilya Kostrikov, Arthur Szlam, and Rob Fergus · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Theophane Weber, Sébastien Racanière, David P. Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, Razvan Pascanu, Peter Battaglia, David Silver, and Daan Wierstra · 2017
Later among the works it cites.
Forward-backward reinforcement learning
Ashley D Edwards, Laura Downs, and James C Davidson · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Closest in time.