Fetching the paper…
Reading the bibliography…
Inverse Reinforcement Learning addresses the problem of inferring an expert's reward function from demonstrations.
Perturbation theory for pseudo-inverses
Perke Wedin · 1973
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
Convergence of a block coordinate descent method for nondifferentiable minimization
Paul Tseng · 2001
Earlier work this paper cites.
Statistical inference
George Casella and Roger L Berger · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Policy gradient methods for robotics
Jan Peters and Stefan Schaal · 2006
Earlier work this paper cites.
An application of reinforcement learning to aerobatic helicopter flight
Pieter Abbeel, Adam Coates, Morgan Quigley, and Andrew Y Ng · 2007
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew L. Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
Navigate like a cabbie: Probabilistic reasoning from observed context-aware behavior
Brian D Ziebart, Andrew L Maas, Anind K Dey, and J Andrew Bagnell · 2008
Cited alongside, same era.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Cited alongside, same era.
A mobile robot that understands pedestrian spatial behaviors
Shu-Yun Chung and Han-Pang Huang · 2010
Cited alongside, same era.
Parametric estimation. finite sample theory
Vladimir Spokoiny et al · 2012
Cited alongside, same era.
Improving hybrid vehicle fuel efficiency using inverse reinforcement learning
Adam Vogel, Deepak Ramachandran, Rakesh Gupta, and Antoine Raux · 2012
Cited alongside, same era.
Noisy and missing data regression: Distribution-oblivious support recovery
Yudong Chen and Constantine Caramanis · 2013
Compatible reward inverse reinforcement learning
Alberto Maria Metelli, Matteo Pirotta, and Marcello Restelli · 2017
Later among the works it cites.
Equivalence between policy gradients and soft q-learning
John Schulman, Xi Chen, and Pieter Abbeel · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adaptive step-size for policy gradient methods
Matteo Pirotta, Marcello Restelli, and Luca Bascetta · 2013
Cited alongside, same era.
Fast and robust least squares estimation in corrupted linear models
Brian McWilliams, Gabriel Krummenacher, Mario Lucic, and Joachim M Buhmann · 2014
Cited alongside, same era.
Reinforcement learning and human behavior
Hanan Shteingart and Yonatan Loewenstein · 2014
Cited alongside, same era.
High-dimensional statistics. spring 2015
Philippe Rigollet · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Inverse reinforcement learning through policy gradient minimization
Matteo Pirotta and Marcello Restelli · 2016
Cited alongside, same era.
Later among the works it cites.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J. Andrew Bagnell, Pieter Abbeel, and Jan Peters · 2018
Later among the works it cites.
Machine theory of mind
Neil Rabinowitz, Frank Perbet, Francis Song, Chiyuan Zhang, SM Ali Eslami, and Matthew Botvinick · 2018
Later among the works it cites.
On-policy robot imitation learning from a converging supervisor
Ashwin Balakrishna, Brijen Thananjeyan, Jonathan Lee, Felix Li, Arsh Zahed, Joseph E. Gonzalez, and Ken Goldberg · 2019
Later among the works it cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Later among the works it cites.
Inverse reinforcement learning with multiple ranked experts
Pablo Samuel Castro, Shijian Li, and Daqing Zhang · 2019
Later among the works it cites.
A first-order approach to accelerated value iteration
Vineet Goyal and Julien Grand-Clement · 2019
Later among the works it cites.
Learning from a learner
Alexis Jacq, Matthieu Geist, Ana Paiva, and Olivier Pietquin · 2019
Later among the works it cites.
Theory of minds: Understanding behavior in groups through inverse planning
Michael Shum, Max Kleiman-Weiner, Michael L Littman, and Joshua B Tenenbaum · 2019
Later among the works it cites.