Fetching the paper…
Reading the bibliography…
Given a dataset of expert demonstrations, inverse reinforcement learning (IRL) aims to recover a reward for which the expert is optimal.
When is a linear control system optimal?
R. E. Kálmán · 1964
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Do people behave according to bellman’s principle of optimality?
J. Rust · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
A. Müller · 1997
Earlier work this paper cites.
Learning agents for uncertain environments
S. Russell · 1998
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
U. Syed and R. E. Schapire · 2007
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Learning robust rewards with adverserial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
A theory of regularized markov decision processes
M. Geist, B. Scherrer, and O. Pietquin · 2019
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans · 2020
Cited alongside, same era.
Policy mirror descent for reinforcement learning: linear convergence, new sampling complexity, and generalized problem classes
G. Lan · 2021
Later among the works it cites.
Online apprenticeship learning
L. Shani, T. Zahavy, and S. Mannor · 2021
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
S. Cen, C. Cheng, Y. Chen, Y. Wei, and Y. Chi · 2022
Later among the works it cites.
Learning from demonstration: Provably efficient adversarial policy imitation with linear function approximation
Z. Liu, Y. Zhang, Z. Fu, Z. Yang, and Z. Wang · 2022
Later among the works it cites.
Proximal point imitation learning
L. Viano, A. Kamoutsi, G. Neu, I. Krawczuk, and V. Cevher · 2022
Later among the works it cites.
A dual approach to constrained markov decision processes with entropy regularization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2021
Cited alongside, same era.
Logistic q-learning
J. Bas-Serrano, S. Curi, A. Krause, and G. Neu · 2021
Cited alongside, same era.
D. Ying, Y. Ding, and J. Lavaei · 2022
Later among the works it cites.
Maximum-likelihood inverse reinforcement learning with finite-time guarantees
S. Zeng, C. Li, A. Garcia, and M. Hong · 2022
Later among the works it cites.
Identifiability and generalizability in constrained inverse reinforcement learning
A. Schlaginhaufen and M. Kamgarpour · 2023
Later among the works it cites.