Fetching the paper…
Reading the bibliography…
Bayesian inference over the reward presents an ideal solution to the ill-posed nature of the inverse reinforcement learning problem.
Minima of functions of several variables with inequalities as side constraints
William Karush · 1939
Earlier work this paper cites.
Nonlinear programming
HW Kuhn and AW Tucker · 1951
Earlier work this paper cites.
Time series analysis of physiological response during icu visitation
Joseph T Hepworth, Sherry Garrett Hendrickson, and Jean Lopez · 1994
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1995
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Markov chain Monte Carlo: stochastic simulation for Bayesian inference
Dani Gamerman and Hedibert F Lopes · 2006
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Map inference for bayesian inverse reinforcement learning
Jaedeug Choi and Kee-Eung Kim · 2011
Earlier work this paper cites.
Bayesian multitask inverse reinforcement learning
Christos Dimitrakakis and Constantin A Rothkopf · 2011
Earlier work this paper cites.
Batch, off-policy and model-free apprenticeship learning
Edouard Klein, Matthieu Geist, and Olivier Pietquin · 2011
Earlier work this paper cites.
Nonlinear inverse reinforcement learning with gaussian processes
Sergey Levine, Zoran Popovic, and Vladlen Koltun · 2011
Earlier work this paper cites.
Preference elicitation and inverse reinforcement learning
Constantin A Rothkopf and Christos Dimitrakakis · 2011
Earlier work this paper cites.
Nonparametric bayesian inverse reinforcement learning for multiple reward functions
Jaedeug Choi and Kee-Eung Kim · 2012
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Infinite time horizon maximum causal entropy inverse reinforcement learning
Michael Bloem and Nicholas Bambos · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Learning robust rewards with adverserial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Later among the works it cites.
Stable baselines
Ashley Hill, Antonin Raffin, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, Rene Traore, Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu · 2018
Later among the works it cites.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Later among the works it cites.
Sören Mindermann, Rohin Shah, Adam Gleave, and Dylan Hadfield-Menell · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Boosted and reward-regularized classification for apprenticeship learning
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2014
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Inverse reinforcement learning with simultaneous estimation of rewards and dynamics
Michael Herman, Tobias Gindele, Jörg Wagner, Felix Schmitt, and Wolfram Burgard · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-Wei, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Cited alongside, same era.
Variational inference: A review for statisticians
David M Blei, Alp Kucukelbir, and Jon D McAuliffe · 2017
Cited alongside, same era.
Later among the works it cites.
Rl baselines zoo
Antonin Raffin · 2018
Later among the works it cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2018
Later among the works it cites.
Deep bayesian reward learning from preferences
Daniel S Brown and Scott Niekum · 2019
Later among the works it cites.
Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, Joern-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky · 2019
Later among the works it cites.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2019
Later among the works it cites.
Truly batch apprenticeship learning with deep successor features
Donghun Lee, Srivatsan Srinivasan, and Finale Doshi-Velez · 2019
Later among the works it cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Siddharth Reddy, Anca D Dragan, and Sergey Levine · 2019
Later among the works it cites.
Bayesian robust optimization for imitation learning
Daniel Brown, Scott Niekum, and Marek Petrik · 2020
Later among the works it cites.
Strictly batch imitation learning by energy-based distribution matching
Daniel Jarrett, Ioana Bica, and Mihaela van der Schaar · 2020
Later among the works it cites.