Fetching the paper…
Reading the bibliography…
The gloabal objective of inverse Reinforcement Learning (IRL) is to estimate the unknown cost function of some MDP base on observed trajectories generated by (approximate) optimal policies.
Muybridges Complete human and Animal locomotion: all 781 plates from the 1887 Animal locomotion
E. Muybridge · 1979
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Escape, avoidance, and imitation: A neural network approach
N. A. Schmajuk and B. S. Zanutto · 1997
Earlier work this paper cites.
Learning agents for uncertain environments
S. Russell · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
A real-world rational agent: unifying old and new ai
P. F. Verschure and P. Althaus · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
An iterative optimal control and estimation design for nonlinear stochastic system
W. Li and E. Todorov · 2006
Earlier work this paper cites.
Maximum margin planning
N. D. Ratliff, J. A. Bagnell, and M. A. Zinkevich · 2006
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
U. Syed and R. E. Schapire · 2007
Earlier work this paper cites.
Apprenticeship learning using linear programming
U. Syed, M. Bowling, and R. E. Schapire · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Using optimization to create self-stable human-like running
K. Mombaur · 2009
Earlier work this paper cites.
Modeling and optimal control of human-like running
G. Schultz and K. Mombaur · 2009
Cited alongside, same era.
Two-factor theory, the actor-critic model, and conditioned avoidance
T. V. Maia · 2010
Cited alongside, same era.
Modeling interaction via the principle of maximum causal entropy
B. D. Ziebart, J. A. Bagnell, and A. K. Dey · 2010
Cited alongside, same era.
Inverse reinforcement learning to control a robotic arm using a brain-computer interface
L. Bougrain, M. Duvinage, and E. Klein · 2012
Cited alongside, same era.
Model-free reinforcement learning with continuous action in practice
T. Degris, P. M. Pilarski, and R. S. Sutton · 2012
Cited alongside, same era.
A kernel two-sample test
A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola · 2012
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Later among the works it cites.
Learning to drive using inverse reinforcement learning and deep q-networks
S. Sharifzadeh, I. Chiotellis, R. Triebel, and D. Cremers · 2016
Later among the works it cites.
Learning robust rewards with adversarial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convex analysis and minimization algorithms I: Fundamentals
J.-B. Hiriart-Urruty and C. Lemaréchal · 2013
Cited alongside, same era.
Generative adversarial networks
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Conditional generative adversarial nets
M. Mirza and S. Osindero · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
The why, what, where, when and how of goal-directed choice: neuronal and computational principles
P. F. Verschure, C. M. Pennartz, and G. Pezzulo · 2014
Cited alongside, same era.
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research, 2018
M. Plappert, M. Andrychowicz, A. Ray, B. McGrew, B. Baker, G. Powell, J. Schneider, J. Tobin, M. Chociej, P. Welinder, V. Kumar, and W. Zaremba · 2018
Later among the works it cites.
Adversarial imitation via variational inverse reinforcement learning
A. H. Qureshi, B. Boots, and M. C. Yip · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
A theory of regularized markov decision processes
M. Geist, B. Scherrer, and O. Pietquin · 2019
Later among the works it cites.
Markov decision process for mooc users behavioral inference
F. Jarboui, C. Gruson-Daniel, A. Durmus, V. Rocchisani, S.-H. G. Ebongue, A. Depoux, W. Kirschenmann, and V. Perchet · 2019
Later among the works it cites.
Situated gail: Multitask imitation using task-conditioned adversarial inverse reinforcement learning
K. Kobayashi, T. Horii, R. Iwaki, Y. Nagai, and M. Asada · 2019
Later among the works it cites.
Ray interference: a source of plateaus in deep reinforcement learning
T. Schaul, D. Borsa, J. Modayil, and R. Pascanu · 2019
Later among the works it cites.
Regularized inverse reinforcement learning
W. Jeon, C.-Y. Su, P. Barde, T. Doan, D. Nowrouzezahrai, and J. Pineau · 2020
Later among the works it cites.
Using inverse reinforcement learning with real trajectories to get more trustworthy pedestrian simulations
F. Martinez-Gil, M. Lozano, I. García-Fernández, P. Romero, D. Serra, and R. Sebastián · 2020
Later among the works it cites.