Fetching the paper…
Reading the bibliography…
Imitation learning algorithms learn viable policies by imitating an expert's behavior when reward signals are not available.
Risk-sensitive markov decision processes
Ronald A Howard and James E Matheson · 1972
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1989
Earlier work this paper cites.
Consideration of risk in reinforcement learning
Matthias Heger · 1994
Earlier work this paper cites.
Robot learning from demonstration
Christopher G Atkeson and Stefan Schaal · 1997
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal · 1997
Earlier work this paper cites.
Neural Networks: A Comprehensive Foundation
Simon Haykin · 1998
Earlier work this paper cites.
Learning agents for uncertain environments
Stuart Russell · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R.S. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Optimization of conditional value-at-risk
R Tyrrell Rockafellar and Stanislav Uryasev · 2000
Earlier work this paper cites.
Q-learning for risk-sensitive control
Vivek S Borkar · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Risk-sensitive reinforcement learning
Oliver Mihatsch and Ralph Neuneier · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Neural systems responding to degrees of uncertainty in human decision-making
Ming Hsu, Meghana Bhatt, Ralph Adolphs, Daniel Tranel, and Colin F Camerer · 2005
Earlier work this paper cites.
An application of reinforcement learning to aerobatic helicopter flight
Pieter Abbeel, Adam Coates, Morgan Quigley, and Andrew Y Ng · 2007
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D. Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Search-based structured prediction
Hal Daumé, John Langford, and Daniel Marcu · 2009
Cited alongside, same era.
Risk-sensitive optimal feedback control accounts for sensorimotor behavior under uncertainty
Arne J Nagengast, Daniel A Braun, and Daniel M Wolpert · 2010
Cited alongside, same era.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Cited alongside, same era.
Risk-averse dynamic programming for markov decision processes
Andrzej Ruszczyński · 2010
Cited alongside, same era.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D Ziebart · 2010
Cited alongside, same era.
Inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2011
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel · 2015
Later among the works it cites.
End to end learning for self-driving cars
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Why is CVaR superior to VaR?(c2009)
Nivine Dalleh · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey J Gordon, and Drew Bagnell · 2011
Cited alongside, same era.
Continuous inverse optimal control with locally optimal examples
Sergey Levine and Vladlen Koltun · 2012
Cited alongside, same era.
Neural prediction errors reveal a risk-sensitive reinforcement-learning process in the human brain
Yael Niv, Jeffrey A Edlund, Peter Dayan, and John P O’Doherty · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Neuroeconomics: Decision making and the brain
Paul W Glimcher and Ernst Fehr · 2013
Cited alongside, same era.
Later among the works it cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Later among the works it cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Later among the works it cites.
Epopt: Learning robust neural network policies using model ensembles
Aravind Rajeswaran, Sarvjeet Ghotra, Sergey Levine, and Balaraman Ravindran · 2016
Later among the works it cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Explaining how a deep neural network trained with end-to-end learning steers a car
Mariusz Bojarski, Philip Yeres, Anna Choromanska, Krzysztof Choromanski, Bernhard Firner, Lawrence Jackel, and Urs Muller · 2017
Closest in time.
Definition of tail risk
Investopedia · 2017
Closest in time.
Risk-sensitive inverse reinforcement learning via coherent risk models
Anirudha Majumdar, Sumeet Singh, Ajay Mandlekar, and Marco Pavone · 2017
Closest in time.
Imitation learning github repository
OpenAI-GAIL · 2017
Closest in time.
On a formal model of safe and scalable self-driving cars
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2017
Closest in time.
Third-person imitation learning
Bradly C Stadie, Pieter Abbeel, and Ilya Sutskever · 2017
Closest in time.