Fetching the paper…
Reading the bibliography…
Bayesian inverse reinforcement learning (IRL) methods are ideal for safe imitation learning, as they allow a learning agent to reason about reward uncertainty and the safety of a learned policy.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Dean A Pomerleau · 1991
Earlier work this paper cites.
Value at risk
Philippe Jorion · 1997
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E Taylor and Peter Stone · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R Duncan Luce · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Risk-sensitive and robust decision-making: a cvar optimization approach
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Optimizing the cvar via sampling
Aviv Tamar, Yonatan Glassner, and Shie Mannor · 2015
Cited alongside, same era.
High-confidence off-policy evaluation
Philip S Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Cited alongside, same era.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Query-efficient imitation learning for end-to-end simulated driving
Jiakai Zhang and Kyunghyun Cho · 2017
Later among the works it cites.
A survey of inverse reinforcement learning: Challenges, methods and progress
Saurabh Arora and Prashant Doshi · 2018
Later among the works it cites.
Playing hard exploration games by watching youtube
Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas · 2018
Later among the works it cites.
Risk-aware active inverse reinforcement learning
Daniel S Brown, Yuchen Cui, and Scott Niekum · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carl Doersch · 2016
Cited alongside, same era.
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, Hado P van Hasselt, and David Silver · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Cited alongside, same era.
Efficient Probabilistic Performance Bounds for Inverse Reinforcement Learning
Daniel S. Brown and Scott Niekum · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Later among the works it cites.
Asking easy questions: A user-friendly approach to active reward learning
Erdem Bıyık, Malayandi Palan, Nicholas C Landolfi, Dylan P Losey, and Dorsa Sadigh · 2019
Closest in time.
Better-than-demonstrator imitation learning via automaticaly-ranked demonstrations
Daniel S Brown, Wonjoon Goo, and Scott Niekum · 2019
Closest in time.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel S. Brown, Wonjoon Goo, Nagarajan Prabhat, and Scott Niekum · 2019
Closest in time.
Importance sampling policy evaluation with an estimated behavior policy
Josiah Hanna, Scott Niekum, and Peter Stone · 2019
Closest in time.
Risk-sensitive generative adversarial imitation learning
Jonathan Lacotte, Mohammad Ghavamzadeh, Yinlam Chow, and Marco Pavone · 2019
Closest in time.
Beyond confidence regions: Tight bayesian ambiguity sets for robust mdps
Marek Petrik and Reazul Hasan Russell · 2019
Closest in time.
Extending deep model predictive control with safety augmented value estimation from demonstrations
Brijen Thananjeyan, Ashwin Balakrishna, Ugo Rosolia, Felix Li, Rowan McAllister, Joseph E Gonzalez, Sergey Levine, Francesco Borrelli, and Ken Goldberg · 2019
Closest in time.