Fetching the paper…
Reading the bibliography…
Consider learning a policy from example expert behavior, without interaction with the expert or access to reinforcement signal.
Neuronlike adaptive elements that can solve difficult learning control problems
A. G. Barto, R. S. Sutton, and C. W. Anderson · 1983
Earlier work this paper cites.
The minimax principle in asymptotic statistical theory
P. W. Millar · 1983
Earlier work this paper cites.
Efficient memory-based learning for robot control
A. W. Moore and T. Hall · 1990
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
D. A. Pomerleau · 1991
Earlier work this paper cites.
Convex Analysis and Minimization Algorithms , volume 305
J.-B. Hiriart-Urruty and C. Lemaréchal · 1996
Earlier work this paper cites.
Learning agents for uncertain environments
S. Russell · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
U. Syed and R. E. Schapire · 2007
Earlier work this paper cites.
Apprenticeship learning using linear programming
U. Syed, M. Bowling, and R. E. Schapire · 2008
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey · 2008
Cited alongside, same era.
On surrogate loss functions and f-divergences
X. Nguyen, M. J. Wainwright, and M. I. Jordan · 2009
Cited alongside, same era.
Learning to search: Functional gradient techniques for imitation learning
N. D. Ratliff, D. Silver, and J. A. Bagnell · 2009
Cited alongside, same era.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Cited alongside, same era.
Modeling interaction via the principle of maximum causal entropy
B. D. Ziebart, J. A. Bagnell, and A. K. Dey · 2010
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Later among the works it cites.
Infinite time horizon maximum causal entropy inverse reinforcement learning
M. Bloem and N. Bambos · 2014
Later among the works it cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Later among the works it cites.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nonlinear inverse reinforcement learning with gaussian processes
S. Levine, Z. Popovic, and V. Koltun · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and D. Bagnell · 2011
Cited alongside, same era.
Elements of information theory
T. M. Cover and J. A. Thomas · 2012
Cited alongside, same era.
Continuous inverse optimal control with locally optimal examples
S. Levine and V. Koltun · 2012
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz
Cited in the paper.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel
Cited in the paper.
Later among the works it cites.
Rlpy: A value-function-based reinforcement learning framework for education and research
A. Geramifard, C. Dann, R. H. Klein, W. Dabney, and J. P. How · 2015
Later among the works it cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Closest in time.
Guided cost learning: Deep inverse optimal control via policy optimization
C. Finn, S. Levine, and P. Abbeel · 2016
Closest in time.
Model-free imitation learning with policy optimization
J. Ho, J. K. Gupta, and S. Ermon · 2016
Closest in time.