Fetching the paper…
Reading the bibliography…
We introduce a novel apprenticeship learning algorithm to learn an expert's underlying reward structure in off-policy model-free \emph{batch} settings.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Maximum margin planning
Nathan D Ratliff, J Andrew Bagnell, and Martin A Zinkevich · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Feature construction for inverse reinforcement learning
Sergey Levine, Zoran Popovic, and Vladlen Koltun · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
In European Workshop on Reinforcement Learning
Edouard Klein, Matthieu Geist, and Olivier Pietquin · 2011
Earlier work this paper cites.
Inverse reinforcement learning through structured classification
Edouard Klein, Matthieu Geist, Bilal Piot, and Olivier Pietquin · 2012
Earlier work this paper cites.
A cascaded supervised learning approach to inverse reinforcement learning
Edouard Klein, Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2013
Earlier work this paper cites.
Learning from demonstrations: Is it worth estimating a reward function?
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2013
Cited alongside, same era.
Boosted and reward-regularized classification for apprenticeship learning
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2014
Cited alongside, same era.
Inverse reinforcement learning via deep gaussian process
Ming Jin, Andreas Damianou, Pieter Abbeel, and Costas Spanos · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Prioritized experience replay
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-wei, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Later among the works it cites.
The third international consensus definitions for sepsis and septic shock (sepsis-3)
Singer Mervyn, Deutschman Clifford S., Seymour Cristopher, and et al · 2016
Later among the works it cites.
Linear feature encoding for reinforcement learning
Zhao Song, Ronald E Parr, Xuejun Liao, and Lawrence Carin · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Later among the works it cites.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, Hado P van Hasselt, and David Silver · 2017
Later among the works it cites.
Learning robust rewards with adversarial inverse reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom Schaul, Antonoglou Ioannis Quan, John, and David Silver · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Inverse reinforcement learning with simultaneous estimation of rewards and dynamics
Michael Herman, Tobias Gindele, Jörg Wagner, Felix Schmitt, and Wolfram Burgard · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Justin Fu, Katie Luo, and Sergey Levine · 2017
Later among the works it cites.
Deep reinforcement learning that matters
Peter Henderson, Risashat Islam, Philip BAchman, Joelle Pineau, Doina Precup, and David Meger · 2017
Later among the works it cites.
Ai safety gridworlds
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg · 2017
Later among the works it cites.
Deep reinforcement learning for sepsis treatment
Aniruddh Raghu, Matthieu Komorowski, Leo Celi Ahmed, Imran, Peter Szolovits, and Marzyeh Ghassemi · 2017
Later among the works it cites.