Fetching the paper…
Reading the bibliography…
Reinforcement learning provides a powerful and general framework for decision making and control, but its application in practice is often hindered by the need for extensive feature and reward engineering.
Learning agents for uncertain environments
Stuart Russell · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Algorithms for reinforcement learning
Andrew Ng and Stuart Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Ng · 2004
Earlier work this paper cites.
Maximum margin planning
Nathan D. Ratliff, J. Andrew Bagnell, and Martin A. Zinkevich · 2006
Earlier work this paper cites.
Boosting structured prediction for imitation learning
Nathan Ratliff, David Bradley, J. Andrew Bagnell, and Joel Chestnutt · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian Ziebart, Andrew Maas, Andrew Bagnell, and Anind Dey · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D. Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian Ziebart · 2010
Cited alongside, same era.
Relative entropy inverse reinforcement learning
Abdeslam Boularias, Jens Kober, and Jan Peters · 2011
Cited alongside, same era.
Nonlinear inverse reinforcement learning with gaussian processes
Sergey Levine, Zoran Popovic, and Vladlen Koltun · 2011
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Later among the works it cites.
Control of memory, active perception, and action in minecraft
Junhyuk Oh, Valliappa Chockalingam, Satinder Singh, and Honglak Lee · 2016
Later among the works it cites.
Repeated inverse reinforcement learning
Kareem Amin, Nan Jiang, and Satinder P. Singh · 2017
Closest in time.
Generalizing skills with semi-supervised reinforcement learning
Chelsea Finn, Tianhe Yu, Justin Fu, Pieter Abbeel, and Sergey Levine · 2017
Closest in time.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Maximum entropy deep inverse reinforcement learning
Markus Wulfmeier, Peter Ondruska, and Ingmar Posner · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine
Cited in the paper.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel
Cited in the paper.
Model-free deep inverse reinforcement learning by logistic regression
Eiji Uchibe · 2017
Closest in time.