Fetching the paper…
Reading the bibliography…
While Reinforcement Learning (RL) aims to train an agent from a reward function in a given environment, Inverse Reinforcement Learning (IRL) seeks to recover the reward function from observing an expert's behavior.
A note on the pure theory of consumer behavior
Samuelson Paul · 1938
Earlier work this paper cites.
Theory of games and economic behavior
J. von Neumann and O. Morgenstern · 1947
Earlier work this paper cites.
Consumption theory in terms of revealed preference
Paul A Samuelson · 1948
Earlier work this paper cites.
Risk aversion in the small and in the large
John W. Pratt · 1964
Earlier work this paper cites.
Aspects of the theory of risk-bearing
Kenneth Joseph Arrow · 1965
Earlier work this paper cites.
A proposal for international monetary reform
James Tobin · 1978
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
Daniel Kahneman and Amos Tversky · 1979
Earlier work this paper cites.
Sample selection bias as a specification error
James J. Heckman · 1979
Earlier work this paper cites.
Large sample properties of generalized method of moments estimators
Lars Peter Hansen · 1982
Earlier work this paper cites.
Finite state markov-chain approximations to univariate and vector autoregressions
George Tauchen · 1986
Earlier work this paper cites.
Optimal replacement of gmc bus engines: An empirical model of harold zurcher
John Rust · 1987
Earlier work this paper cites.
Conditional choice probabilities and the estimation of dynamic models
V. Joseph Hotz and Robert A. Miller · 1993
Earlier work this paper cites.
Identification of causal effects using instrumental variables
Joshua D. Angrist, Guido W. Imbens, and Donald B. Rubin · 1996
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Learning agents for uncertain environments (extended abstract)
Stuart Russell · 1998
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Maximum margin planning
N. D. Ratliff, J. A. Bagnell, and M. A. Zinkevich · 2006
Earlier work this paper cites.
Maximum margin planning
Nathan D Ratliff, J Andrew Bagnell, and Martin A Zinkevich · 2006
Earlier work this paper cites.
Bayesian inverse reinforcement learning
D. Ramachandran and E. Amir · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Apprenticeship learning using linear programming
U. Syed, M. Bowling, and R.E. Schapire · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Inverse optimal control with linearly-solvable MDPs
K. Dvijotham and E. Todorov · 2010
Earlier work this paper cites.
254a, notes 3a: Eigenvalues and sums of hermitian matrices
Terence Tao · 2010
Earlier work this paper cites.
Relative entropy inverse reinforcement learning
Abdeslam Boularias, Jens Kober, and Jan Peters · 2011
Cited alongside, same era.
Nonlinear inverse reinforcement learning with Gaussian processes
S. Levine, Z. Popović, and V. Koltun · 2011
Cited alongside, same era.
Dimension-free tail inequalities for sums of random matrices, 2011
Daniel Hsu, Sham M. Kakade, and Tong Zhang · 2011
Cited alongside, same era.
On the sample complexity of reinforcement learning with a generative model, 2012
Mohammad Gheshlaghi Azar, Remi Munos, and Bert Kappen · 2012
Cited alongside, same era.
Random design analysis of ridge regression
Daniel Hsu, Sham M. Kakade, and Tong Zhang · 2012
Cited alongside, same era.
Dynamic models and structural estimation in corporate finance
Ilya A Strebulaev and Toni M Whited · 2012
Cited alongside, same era.
Action robust reinforcement learning and applications in continuous control
Chen Tessler, Yonathan Efroni, and Shie Mannor · 2019
Later among the works it cites.
A Theory of Regularized Markov Decision Processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Later among the works it cites.
Reinforcement learning in economics and finance, 2020
Arthur Charpentier, Romuald Elie, and Carl Remlinger · 2020
Later among the works it cites.
Quantifying differences in reward functions, 2020
Adam Gleave, Michael Dennis, Shane Legg, Stuart Russell, and Jan Leike · 2020
Later among the works it cites.
Efficient exploration of reward functions in inverse reinforcement learning via bayesian optimization, 2020
Sreejith Balakrishnan, Quoc Phong Nguyen, Bryan Kian Hsiang Low, and Harold Soh · 2020
Later among the works it cites.
Safe imitation learning via fast bayesian reward inference from preferences, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning from demonstrations: Is it worth estimating a reward function?
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2013
Cited alongside, same era.
Solving relational mdps with exogenous events and additive rewards, 2013
S. Joshi, R. Khardon, P. Tadepalli, A. Raghavan, and A. Fern · 2013
Cited alongside, same era.
Bridging the gap between imitation learning and inverse reinforcement learning
Bilal Piot, Matthieu Geist, and Olivier Pietquin · 2016
Cited alongside, same era.
Model-free imitation learning with policy optimization
J. Ho, J. K. Gupta, and S. Ermon · 2016
Cited alongside, same era.
Towards resolving unidentifiability in inverse reinforcement learning, 2016
Kareem Amin and Satinder Singh · 2016
Cited alongside, same era.
Inverse reinforcement learning through policy gradient minimization
Matteo Pirotta and Marcello Restelli · 2016
Cited alongside, same era.
Daniel S. Brown, Russell Coleman, Ravi Srinivasan, and Scott Niekum · 2020
Later among the works it cites.
Inverse reinforcement learning from a gradient-based learner
Giorgia Ramponi, Gianluca Drappo, and Marcello Restelli · 2020
Later among the works it cites.
State-only imitation with transition dynamics mismatch
Tanmay Gangwani and Jian Peng · 2020
Later among the works it cites.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
Daniel S. Brown, Wonjoon Goo, and Scott Niekum · 2020
Later among the works it cites.
Interpretable batch irl to extract clinician goals in icu hypotension management
Srivatsan Srinivasan and Finale Doshi-Velez · 2020
Later among the works it cites.
New developments in revealed preference theory: decisions under risk, uncertainty, and intertemporal choice
Federico Echenique · 2020
Later among the works it cites.
Robust reinforcement learning via adversarial training with langevin dynamics
Parameswaran Kamalaruban, Yu-Ting Huang, Ya-Ping Hsieh, Paul Rolland, Cheng Shi, and Volkan Cevher · 2020
Later among the works it cites.
Identifiability in inverse reinforcement learning
Haoyang Cao, Samuel Cohen, and Lukasz Szpruch · 2021
Later among the works it cites.
Reward (mis)design for autonomous driving, 2021
W. Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt, and Peter Stone · 2021
Later among the works it cites.
Dealing with multiple experts and non-stationarity in inverse reinforcement learning: an application to real-life problems
Amarildo Likmeta, Alberto Maria Metelli, Giorgia Ramponi, Andrea Tirinzoni, Matteo Giuliani, and Marcello Restelli · 2021
Later among the works it cites.
Explicable reward design for reinforcement learning agents
Rati Devidze, Goran Radanovic, Parameswaran Kamalaruban, and Adish Singla · 2021
Later among the works it cites.
Strictly batch imitation learning by energy-based distribution matching, 2021
Daniel Jarrett, Ioana Bica, and Mihaela van der Schaar · 2021
Later among the works it cites.
Value alignment verification
Daniel S Brown, Jordan Schneider, Anca Dragan, and Scott Niekum · 2021
Later among the works it cites.
Reward identification in inverse reinforcement learning
Kuno Kim, Shivam Garg, Kirankumar Shiragur, and Stefano Ermon · 2021
Later among the works it cites.
Provably efficient learning of transferable rewards
Alberto Maria Metelli, Giorgia Ramponi, Alessandro Concetti, and Marcello Restelli · 2021
Later among the works it cites.
Inverse decision modeling: Learning interpretable representations of behavior
Daniel Jarrett, Alihan Hüyük, and Mihaela Van Der Schaar · 2021
Later among the works it cites.
On the expressivity of markov reward
David Abel, Will Dabney, Anna Harutyunyan, Mark K Ho, Michael Littman, Doina Precup, and Satinder Singh · 2021
Later among the works it cites.
Exploiting action impact regularity and exogenous state variables for offline reinforcement learning, 2021
Vincent Liu, James Wright, and Martha White · 2021
Later among the works it cites.
Robust inverse reinforcement learning under transition dynamics mismatch
Luca Viano, Yu-Ting Huang, Parameswaran Kamalaruban, Adrian Weller, and Volkan Cevher · 2021
Later among the works it cites.
Invariance in policy optimisation and partial identifiability in reward learning, 2022
Joar Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave · 2022
Closest in time.
Option compatible reward inverse reinforcement learning
Rakhoon Hwang, Hanjin Lee, and Hyung Ju Hwang · 2022
Closest in time.
Robust learning from observation with model misspecification
Luca Viano, Yu-Ting Huang, Parameswaran Kamalaruban, Craig Innes, Subramanian Ramamoorthy, and Adrian Weller · 2022
Closest in time.