Fetching the paper…
Reading the bibliography…
Inverse Reinforcement Learning (IRL) algorithms infer a reward function that explains demonstrations provided by an expert acting in the environment.
Information-theoretical optimization techniques
Flemming Topsøe · 1979
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart J Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Brian D. Ziebart, J. Andrew Bagnell, and Anind K. Dey · 2010
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Maximum entropy deep inverse reinforcement learning
Markus Wulfmeier, Peter Ondruska, and Ingmar Posner · 2015
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Cited alongside, same era.
Repeated inverse reinforcement learning
Kareem Amin, Nan Jiang, and Satinder Singh · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D. Dragan, S. Shankar Sastry, and Sanjit A. Seshia · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Cited alongside, same era.
A practical approach to insertion with variable socket position using deep reinforcement learning
Mel Vecerik, Oleg Sushkov, David Barker, Thomas Rothörl, Todd Hester, and Jon Scholz · 2019
Later among the works it cites.
seals: Suite of environments for algorithms that learn specifications
Adam Gleave, Pedro Freire, Steven Wang, and Sam Toyer · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca Dragan · 2020
Later among the works it cites.
Keynote talk
Andrej Karpathy · 2020
Later among the works it cites.
Upgrading Autopilot: Seeing the world in radar
Tesla · 2020
Later among the works it cites.
Identifiability in inverse reinforcement learning
Haoyang Cao, Samuel N. Cohen, and Lukasz Szpruch · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Inverse reinforcement learning for video games
Aaron Tucker, Adam Gleave, and Stuart Russell · 2018
Cited alongside, same era.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson · 2019
Cited alongside, same era.
Preferences implicit in the state of the world
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander, Pieter Abbeel, and Anca Dragan · 2019
Cited alongside, same era.
A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine
Cited in the paper.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel
Cited in the paper.
Quantifying differences in reward functions
Adam Gleave, Michael Dennis, Shane Legg, Stuart Russell, and Jan Leike · 2021
Later among the works it cites.
Reward identification in inverse reinforcement learning
Kuno Kim, Shivam Garg, Kirankumar Shiragur, and Stefano Ermon · 2021
Later among the works it cites.
Invariance in policy optimisation and partial identifiability in reward learning
Joar Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave · 2022
Closest in time.