Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1989
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 1994
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1995
Earlier work this paper cites.
The maxq method for hierarchical reinforcement learning
Thomas G Dietterich et al · 1998
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart Russell · 1998
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
Stefan Schaal · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
State aggregation in markov decision processes
Zhiyuan Ren and Bruce H Krogh · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi, Benjamin Recht, et al · 2007
Earlier work this paper cites.
Using bisimulation for policy transfer in mdps
Pablo Castro and Doina Precup · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
On the convergence of the empirical distribution
Original
Daniel Berend and Aryeh Kontorovich · 2012
Earlier work this paper cites.
Variational intrinsic control
Original
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Earlier work this paper cites.